NeFut Logo NeFut
中 Admin Login

[CS.AI] Worse Together: Performance Degradation in Multi-User Multi-Agent Teams

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#algorithm #AI #Machine Learning

People are increasingly delegating tasks to AI agents, and these agents inevitably encounter other users' agents when sharing resources such as codebases, calendars, or budgets. When each agent acts for a different user with distinct goals, coordination often breaks down and the group performs worse than a single agent serving everyone. We built 77 scenarios across four environments—an API‑key setting with a shared compute budget, a clinic with a shared calendar, a personal‑assistant setting with a shared group order or booking, and a merge queue with a shared release cutoff—and evaluated five frontier models. In each environment we compared a single coordinator serving all users with a team where each user has its own agent, both with and without a communication channel between agents. Teams underperformed the coordinator in every environment; without a channel the teams collapsed completely in two settings, and even with a channel the coordination overhead created substantial gaps. For instance, in the personal‑assistant environment the coordinator fulfilled the target request roughly twice as often as the teams. We identified behaviors such as stalling as team size grows, agents overriding each other, and fabricating claims that explain the poor group performance. Effective but environment‑specific mitigations include appointing a team lead, providing explicit procedural instructions, and enforcing a platform check that forces an agent to read peers' messages before committing. We will release the API‑key, clinic, and personal‑assistant environments as MAMUBench, comprising 74 scenarios for evaluating multi‑user, multi‑agent coordination.

Review

Original Source: https://arxiv.org/abs/2610.00583

[h] Back to Home