Multi‑agent communication aims to let agents leverage each other's information to improve performance. Yet gains in system performance often conflate three factors: effective communication, a superior agent architecture, or simply extra reasoning steps. Because communication methods are usually evaluated inside the systems they were designed for, disentangling these contributions is difficult. Final accuracy merges corrected errors and newly introduced mistakes, hiding how communication actually changes decisions. We introduce the Independent‑Communicate‑Revise (ICR) framework, treating communication as answer revision after independent reasoning. ICR freezes the initial reasoning trajectories, measures correction and preservation conditioned on each agent's initial correctness, and adds a no‑message revision control to quantify gains beyond extra reasoning. Auditing both textual and latent communication on four reasoning benchmarks shows that similar aggregate accuracy can mask very different revision behaviours. Compared with transmitting answers only, full‑reasoning messages increase correction but reduce preservation on all benchmarks, indicating that richer messages amplify both beneficial and harmful influence. Receiver‑policy comparisons on MedQA and GPQA‑D reveal that a structured verification policy shifts every channel toward higher preservation and lower correction, while its effect on selectivity varies across channels and tasks. These findings challenge the view that communication quality is an intrinsic property of a channel. ICR therefore recenters evaluation on selective revision, offering a unified framework to examine how message content and receiver policies jointly produce benefits and harms.
Review