Stochastic Decision Horizons for Survival-Based Constrained Reinforcement Learning
We propose stochastic decision horizons (SDH), a theoretically grounded framework for survival-based constrained reinforcement learning from per-step violation signals. Rather than treating violations as expenditure of a cumulative budget, SDH shortens the effective decision horizon when violations occur, reducing current reward credit and future bootstrapping. We show that this construction represents a geometric-horizon survival constraint and, for a fixed Lagrange multiplier, remains the same SDH problem up to a constant reward offset that acts as an implicit incentive for survival. Using control as inference, we derive two off-policy and replay-compatible, regularized algorithms with post-violation semantics. Absorbing-state ends the decision process, so only surviving decisions pay policy KL, yielding AS-SAC. Virtual termination stops reward credit while the process continues, retaining ordinary discounted KL and yielding VT-MPO. To connect SDH with cumulative-cost CMDPs, we introduce violation-depth profiles. Under exponential-tail single-scale profiles, the SDH statistic determines cumulative cost; departures from this form mark where the connection breaks. Experiments validate both the method and this predicted scope. On the 90-muscle H2190 humanoid in Hyfydy, VT-MPO matches the EWA reference in peak gait realism, with its highest observed gait-match checkpoint at 21M environment steps versus 90M for EWA and markedly more stable training. On Safety Gymnasium, violation-depth profiles predict the regimes in which SDH achieves strong reward-violation trade-offs.