Search arXiv⌕ Search

arXiv subjects

Jhojan A. Rodriguez-Gil

Publications and source records attributed to Jhojan A. Rodriguez-Gil.

1 recordsLinked to original sources

Training Dynamics and Induced Preconditioning of ReLU Policy Gradient for Scalar LQR

Policy gradient is not invariant to reparameterization; the training dynamics of the weights of a neural controller differ from those of the induced gains. We make that difference exact for deterministic scalar discounted linear-quadratic regulation with bias-free one-hidden-layer ReLU policies. We show that policy gradient on the weights induces preconditioning on the gains' dynamics, alongside an explicit quadratic finite-step correction and an activation-group error. Moreover, we identify explicit conditions on the step sizes and initialization parameters that guarantee that the sequence of generated controllers remains stabilizing throughout training and converges to the optimal gains with high probability. For a given system and cost parameters, the sufficient step-size conditions imply contraction bounds independent of the network width. Numerical analysis illustrates how the neural realization changes the induced gain dynamics while maintaining stability across training.

math.OC↗