arXiv · 2601.19499
Reversible Simplex Supervision with Post-Action Debt Accounting for Goal-Reaching RL
Abstract
Deploying reinforcement learning (RL) on multi-tonne robots calls for supervisory mechanisms that address both operational safety and progress toward task completion. However, repeated switching need not preserve task progress when a learned action increases storage before recovery takes control. We introduce reversible Simplex supervision with post-action debt accounting for a frozen finite-state policy and robust-adaptive recovery. Under exact sampled-state information and stated model and certificate conditions, we prove that recovery repayment exceeding a uniform triggering-edge debt bound guarantees finite switching and finite-sample goal entry. We formulate reachability-based certificate constructions for establishing these sufficient conditions. In 20 matched simulations, goal-entry counts are 20 with debt gating and 18 without it; the two remaining runs terminate under the supervisor's admissibility stopping rule. On an experimental 6000 kg robot, 24 asphalt and soft-terrain trials evaluate 50 ms supervision above a 1 kHz actuator stack; all eight triggered recoveries complete debt-gated re-entry. The experiments demonstrate the supervisory mechanism in the tested trials.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mehdi Heydari Shahna, Joongheon Kim, Jouni Mattila. 2026-09-15. Reversible Simplex Supervision with Post-Action Debt Accounting for Goal-Reaching RL. https://arxiv.org/abs/2601.19499
Cite the original work for its findings. Save a collection to share your selection of sources.