arXiv · 2609.21433
Benign Nonconvex Landscape for Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action Spaces
Abstract
We study the optimization landscape for infinite-horizon discounted Markov decision processes (MDPs) with general state and action spaces under structured stationary policy classes. A general weighted policy-iteration approach to establishing global convergence guarantees for policy gradient methods requires closure under weighted policy improvement at every policy, a property that may fail even when the policy class contains an optimal policy. To address this issue, we propose weaker conditions that guarantee the absence of suboptimal stationary points and establish the Polyak--Lojasiewicz--Kurdyka (PLK) condition for the policy gradient objective with a finite concentrability coefficient. We also establish the PLK condition from a policy-improvement bound that holds at every state, without a concentrability assumption. Our general results encompass settings covered by the earlier framework when the common standing assumptions hold for the same policy class and parameter domain. We further verify our proposed conditions for two operations models: inventory systems with Markov-modulated demand and stochastic cash-balance problems. For both models, the Bellman equation yields approximate convexity of the Q-value functions in the action variable, with deviations controlled by the first-order stationarity measure. These estimates establish exponent-one and, under additional curvature assumptions, exponent-two PLK conditions, which, together with Lipschitz continuity of the policy gradient, imply an $\mathcal{O}(1/ε)$ iteration complexity and linear convergence, respectively, for projected gradient descent using exact policy gradients. To the best of our knowledge, we provide the first non-asymptotic convergence rates for solving infinite-horizon discounted inventory systems with Markov-modulated demand and stochastic cash-balance problems using policy gradient methods.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Xin Chen, Minda Zhao. 2026-09-18. Benign Nonconvex Landscape for Policy Optimization: Infinite-Horizon Discounted MDPs with General State and Action Spaces. https://arxiv.org/abs/2609.21433
Cite the original work for its findings. Save a collection to share your selection of sources.