Multi‑agent large language model (LLM) systems are often assumed to improve as the team grows, yet the actual scaling depends on the task structure. This work adopts Steiner's taxonomy of group tasks, focusing on disjunctive and compensatory tasks as a framework for analyzing multi‑agent scaling. Agents are modeled as conditionally independent given an item, leading to large‑team limits: plurality voting converges to the model's modal answer, while averaging converges to the model's item‑level bias.
Experiments across selected benchmarks, 13 open‑weight models, and teams up to 30 agents reveal qualitatively different scaling patterns. For disjunctive tasks, the probability that at least one agent is correct rises by 5–20 points with team size, but direct plurality voting over agents' answers captures almost none of this potential, improving by only about 0.5 points on average. Multi‑round revision boosts accuracy considerably, yet the gain from a single peer is nearly the same as from 29 peers.
In contrast, scaling offers little benefit on Fermi estimation, a naturally aggregable compensatory task. Item‑level biases shared across model samples account for roughly 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but still does not surpass the strongest single model on disjunctive tasks.
These findings demonstrate that task structure together with the mechanism for combining member outputs fundamentally determines team scaling.
Review