We adapted a three‑stage deliberation protocol originally designed for humans to three families of large language models (LLMs) and evaluated it across four domains with increasing real‑world stakes: visual numeric estimation (Study 1), peer review of machine‑learning papers (Study 2), detection of hidden malicious behavior by an AI agent (Study 3), and sports forecasting against a live prediction market (Study 4). In each study, models first produced independent answers, then engaged in multiple rounds of group deliberation, and finally the consensus estimates were averaged. Across all domains, deliberation reduced the collective error beyond what passive aggregation of independent responses achieved, and the post‑deliberation individual judgments retained this collective gain. Crucially, the benefit required model diversity: groups composed of clones of a single model did not show improvement. These findings establish machine deliberation as a general‑purpose aggregation mechanism and highlight diversity as an active ingredient.
Review