arXiv · 2609.34695
Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control
Abstract
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective from numerical regularizers, leading to brittle hyperparameter tuning. To determine which quantities a reward must contain, we study the stabilization control problem with a focus on zeroth-order (configuration) and first-order (velocity) information. We theoretically and empirically demonstrate that policy gradient methods can successfully solve stabilization tasks without first-order reward terms, adding such terms can instead introduce severe sensitivity as their scale grows. Conversely, our findings confirm that reward functions must be zeroth-order complete over goal-relevant coordinates, while the first-order state remains necessary in the policy observation under our low-dissipation assumptions. Overall, these results provide actionable and principled guidance for reward design in robotic RL.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yisheng Zhang, Tao Wang, Sicun Gao. 2026-09-28. Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control. https://arxiv.org/abs/2609.34695
Cite the original work for its findings. Save a collection to share your selection of sources.