This paper investigates whether the gains of Multi‑Agent Debate (MAD) stem from cognitive diversity among agents. We conduct measurable experiments on small open‑weight language models, selecting 23 models from 11 vendors, covering five tasks and running over 5,500 debate and control trials.\ \ Diversity is introduced along three axes:\
- Persona prompting\
- Sampling temperature\
- Model identity\ Each debate configuration is paired with a majority‑vote control that matches the generation budget.\ \ The results reject the “diversity‑driven” hypothesis:\
- On all axes, debate does not achieve a significant additional boost\
- Debate outperforms single‑model inference by 3–7 points (where tasks have headroom), but under equal‑budget conditions it merely ties or falls behind self‑consistency sampling, which costs about 1.6× wall‑clock time and 3.4× token usage\
- Persona prompting reduces accuracy; a dose‑response study over the full combinatorial persona space shows this is a “persona tax” rather than a diversity tax: redundant personas hurt the most, while maximally diverse teams recover only part of the loss\
- Mixed‑model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity\
- Nearly all of debate’s benefit comes from the first exchange of answers\ \ We also identify a pervasive measurement hazard: debate transcripts often overflow the context window and are silently truncated. Correcting this alone shifts the debate‑vs‑sampling comparison from –1.8 points to parity.\ \ In summary, previously reported MAD gains appear to be an ensemble‑sampling effect rather than a product of cognitive diversity. This work establishes a budget‑matched, contamination‑checked baseline that future debate mechanisms must surpass.\ \ Review