arXiv · 2609.27312
Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning
Abstract
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploitable by deliberate attacks. We propose Safety to Competence (S2C), a two-stage RL framework that separates safety synthesis from competitive task learning. We formulate competitive interactions as safety-critical Markov games and prove that perfect filtering preserves policy non-exploitability when all players commit to safe maneuvers. S2C learns a robust safety filter via adversarial RL, embeds it in the environment during task policy training, and retains the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe RL baselines, achieving the highest win rate and Elo rating, and the lowest exploitability. Hardware stress tests against a human opponent confirm S2C's competence.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu. 2026-09-23. Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning. https://arxiv.org/abs/2609.27312
Cite the original work for its findings. Save a collection to share your selection of sources.