Search arXiv⌕ Search

arXiv subjects

Shoju Enami

Publications and source records attributed to Shoju Enami.

4 recordsLinked to original sources

On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems

Mutual information regularization has recently been studied in reinforcement learning (RL) as an extension of entropy regularization, in which not only the policy but also the reference distribution (prior) is optimized, and has also been used for privacy preserving policy design. To obtain theoretical insight into how this joint optimization affects policy stochasticity, we study a model-based linear-quadratic setting that preserves this key structure while permitting an explicit analysis. Specifically, we consider a mutual information optimal control problem (MIOCP) for stochastic discrete-time linear systems with quadratic costs and Gaussian policies and priors. We first review alternating optimization of the policy and the prior with a technical extension. We then establish the existence of an optimal solution, characterize its covariance matrices, and derive sufficient conditions under which the optimal policy is either a stochastic feedback policy or a deterministic open-loop policy. We further interpret this characterization in terms of the state information disclosed through the control input. We also analyze the limiting behavior of alternating optimization under the sufficient conditions. The validity of the theoretical results is demonstrated through numerical experiments.

math.OC↗

Mutual Information Optimal Density Control of Linear Systems and Generalized Schrödinger Bridges with Reference Refinement

We consider a mutual information (MI) regularized version of optimal density control of a discrete-time linear system. MI optimal control has been proposed as an extension of maximum entropy optimal control to trade off between control performance and benefits provided by stochastic inputs. MI regularization induces stochasticity in the policy, which poses challenges for applications of MI optimal control in safety-critical scenarios. To remedy this situation, we impose Gaussian density constraints at specified times to directly control state uncertainty. For this MI optimal density control problem, we propose an alternating optimization algorithm and derive the closed form of each step in the algorithm. In addition, we reveal that the alternating optimization of the MI optimal density control problem coincides with that of the so-called generalized Schrödinger bridge problem associated with the discrete-time linear system.

math.OC↗

Mutual Information Optimal Control of Discrete-Time Linear Systems

In this paper, we formulate a mutual information optimal control problem (MIOCP) for discrete-time linear systems. This problem can be regarded as an extension of a maximum entropy optimal control problem (MEOCP). Differently from the MEOCP where the prior is fixed to the uniform distribution, the MIOCP optimizes the policy and prior simultaneously. As analytical results, under the policy and prior classes consisting of Gaussian distributions, we derive the optimal policy and prior of the MIOCP with the prior and policy fixed, respectively. Using the results, we propose an alternating minimization algorithm for the MIOCP. Through numerical experiments, we discuss how our proposed algorithm works.

math.OC↗

A proposal of adaptive parameter tuning for robust stabilizing control of $N$--level quantum angular momentum systems

Stabilizing control synthesis is one of the central subjects in control theory and engineering, and it always has to deal with unavoidable uncertainties in practice. In this study, we propose an adaptive parameter tuning algorithm for robust stabilizing quantum feedback control of $N$-level quantum angular momentum systems with a robust stabilizing controller proposed by [Liang, Amini, and Mason, SIAM J. Control Optim., 59 (2021), pp. 669-692]. The proposed method ensures local convergence to the target state. Besides, numerical experiments indicate its global convergence if the learning parameters are adequately determined.

math.OC↗