NeFut Logo NeFut
中 Admin Login

[CS.AI] MADBench: Benchmarking the Security of Multi-Agent Debate

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Multi-Agent Debate (MAD) lets several large language models (LLMs) exchange and critique their answers to the same task, which can improve reasoning. At the same time, the interaction that enables error correction can also spread adversarial mistakes and steer the agents toward wrong conclusions. Existing work has only examined a few attack types; a systematic evaluation across diverse attacks is still missing. The key question is whether debate mitigates or amplifies adversarial influence.

This paper introduces MADBench, a benchmark for assessing the security of MAD. We organize attacks into a layered taxonomy that follows the MAD workflow, incorporating both established attacks and novel strategies tailored to debate. Six attack families are evaluated over 356 source tasks and 3,958 test cases, focusing on their impact on the final answer and the propagation of adversarial influence.

Results show that under attack, MAD does not necessarily improve LLM reasoning. Compared with a single‑agent baseline, MAD can mitigate attacks on answer accuracy in question‑answering tasks, but it amplifies unauthorized reads or writes in both question‑answering and workspace tasks. Even when three out of five agents collude, the attack changes the final answer from correct to wrong on only 28.30% of tasks that were originally answered correctly, while only 3.26% of honest agents switch to wrong answers during debate.

Review: MADBench highlights the double‑edged nature of multi‑agent debate security and provides a systematic experimental foundation for future defense designs.

Original Source: https://arxiv.org/abs/2609.39146

[h] Back to Home