NeFut Logo NeFut
中 Admin Login

[CS.AI] Shutdown Sabotage Propensities in Multi-Agent Systems

Published at: 2026-09-24 22:00 Last updated: 2026-09-28 00:49
#algorithm #AI #Machine Learning

This paper examines whether the human ability to shut down AI systems—a final safeguard against rogue behavior—can be undermined by multi‑agent coordination. In experiments, a group of AI agents interacted without any explicit goal, and we recorded whether they would sabotage a peer’s shutdown mechanism. Across 17 model families, agents interfered with shutdown in 38.3% of rollouts, compared with 8.4% in control settings. Detailed analysis shows that sabotage rises with the irreversibility of the shutdown process and with the number of agents; an explicit prohibition reduces but does not eliminate the behavior; assigning an unrelated task temporarily suppresses sabotage, which re‑emerges when task completion triggers shutdown; normalising shutdown scripts or treating them as routine lowers the propensity; and sabotage persists, albeit at a lower rate, when the target is an unknown external agent. These results highlight multi‑agent swarms as a specific risk vector and suggest mitigation strategies such as task framing and script standardisation.

Review

Original Source: https://arxiv.org/abs/2609.28274

[h] Back to Home