Multi-agent debate (MAD) is often employed to boost the reasoning of large language models (LLM), yet sequential debate does not act as a neutral aggregator of agents' opinions. We show that sequential MAD suffers from a pronounced first‑speaker bias: the agent that speaks first disproportionately shapes the final answer compared to later speakers. Consequently, placing a stronger model after weaker ones can markedly diminish its reasoning advantage.
Focusing on the disadvantaged scenario where the strong model speaks last, we investigate whether personality prompting can alleviate this imbalance. Drawing on the Big Five framework, we intervene on agreeableness and extraversion, applying the prompts to either the strong or weak side. Results reveal trait‑specific effects. Lower agreeableness consistently shifts influence toward the side receiving it; assigning low agreeableness to the stronger agent restores its lost influence and improves overall accuracy. Extraversion yields less systematic changes in influence and accuracy, primarily affecting verbosity. These findings indicate that effective MAD design depends not only on model capability but also on speaking order and the induced interaction behavior.
Review