Robots deployed in competitive tasks must outmaneuver opponents while preserving safety. Existing safe reinforcement learning methods train a single policy to achieve success and avoid failures simultaneously, which complicates training and leaves the policy vulnerable to deliberate attacks. We introduce Safety to Competence (S2C), a two‑stage RL framework that separates safety synthesis from competitive task learning. Competitive interactions are modeled as safety‑critical Markov games, and we prove that perfect filtering preserves non‑exploitability when all players commit to safe actions. S2C first learns a robust safety filter via adversarial RL, then embeds this filter into the environment during task‑policy training, retaining the same filter at deployment. In simulated touchdown games, S2C outperforms eight safe‑RL baselines, achieving the highest win rate and Elo rating while dramatically reducing exploitability. Hardware stress tests against a human opponent confirm S2C's competence.
Review