NeFut Logo NeFut
中 Admin Login

[CS.AI] A Ghost in Long-Horizon Agents: Governance Hazard from Overlooked Safety Constraints across Turns

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Long-horizon agents are increasingly tasked with assisting humans on complex problems. Their extended interaction history introduces an under‑explored execution‑safety issue: under seemingly benign conditions an agent may take an action that violates a safety constraint specified many turns earlier. We name this failure mode Governance Hazard from Overlooked Safety Constraints across Turns (GHOST), which can cause irreversible damage. Experiments show that GHOST events are not isolated; on GPT‑5.5 they occur with a rate of 11.5%. Theoretically, if the residual conditional violation hazard for each safe prefix is bounded below by a non‑summable sequence, the execution almost surely enters the hazard region. Leveraging this insight we propose STAR‑Guard, a two‑layer defense that couples historical semantic safety‑constraint restoration with pre‑execution audit. The restoration layer re‑activates applicable constraints to reduce unsafe proposals, while the deterministic audit layer prevents any remaining violations from reaching the environment. Consistent with this design, experiments under the GPT‑5.5 setup observe zero GHOST events.

Review: GHOST highlights a temporal blind spot in safety management for long‑term agents, and STAR‑Guard’s layered approach offers a practical mitigation strategy that could shape future safe‑agent development.

Original Source: https://arxiv.org/abs/2610.02664

[h] Back to Home