NeFut Logo NeFut
中 Admin Login

[CS.AI] SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

Published at: 2026-08-21 11:26 Last updated: 2026-08-21 22:00
#algorithm #Machine Learning #Artificial Intelligence

Vision‑language‑model based embodied agents can follow instructions but often breach safety constraints during execution, a problem framed as interactive safety. Safety and task success are distinct objectives, and safety violations occur only at a few critical steps in a trajectory, making standard supervision insufficient. Imitating safe trajectories teaches the behavior without explaining why it is safe, while contrasting arbitrary safe and unsafe trajectories mixes the safety signal with unrelated differences.

SafeBranch addresses this by constructing branch‑pairs from the agent’s own unsafe rollouts via environment rollback. For each unsafe rollout the system rolls back to the safety‑critical step that caused the violation, queries the agent for a safe alternative action at that state, and pairs the original action with the alternative so that the two branches differ only at that step. After training, the actor can act safely at deployment without any critic in the loop.

Evaluations on IS‑Bench, SafetyALFRED and out‑of‑distribution variants with unseen tasks and objects show that SafeBranch maintains task success while dramatically improving safety, achieving roughly ten times more safe successes than an untrained baseline on the unseen‑object variant.

Blogger's Review: The rollback‑pairing mechanism elegantly turns sparse safety signals into dense, learnable contrastive examples, eliminating the need for an external safety critic while preserving performance. Its strong results on diverse benchmarks suggest strong potential for more complex embodied environments.

Original Source: https://arxiv.org/abs/2608.19729

[h] Back to Home