arXiv · 2507.21543
On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems
Abstract
Mutual information regularization has recently been studied in reinforcement learning (RL) as an extension of entropy regularization, in which not only the policy but also the reference distribution (prior) is optimized, and has also been used for privacy preserving policy design. To obtain theoretical insight into how this joint optimization affects policy stochasticity, we study a model-based linear-quadratic setting that preserves this key structure while permitting an explicit analysis. Specifically, we consider a mutual information optimal control problem (MIOCP) for stochastic discrete-time linear systems with quadratic costs and Gaussian policies and priors. We first review alternating optimization of the policy and the prior with a technical extension. We then establish the existence of an optimal solution, characterize its covariance matrices, and derive sufficient conditions under which the optimal policy is either a stochastic feedback policy or a deterministic open-loop policy. We further interpret this characterization in terms of the state information disclosed through the control input. We also analyze the limiting behavior of alternating optimization under the sufficient conditions. The validity of the theoretical results is demonstrated through numerical experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shoju Enami, Kenji Kashima. 2026-09-19. On Policy Stochasticity in Mutual Information Optimal Control of Linear Systems. https://arxiv.org/abs/2507.21543
Cite the original work for its findings. Save a collection to share your selection of sources.