This work evaluates how replacing a single LLM agent with a collaborating team affects accuracy across eight orchestration architectures. Five instruction‑tuned 7‑9B models were tested on five short‑answer benchmarks and an executable‑code benchmark, with call budgets up to 30. The returns to scaling are sharply task‑dependent: on the two arithmetic word‑problem suites (GSM8K, GSMHard) accuracy rises up to 17 points when calls increase from three to thirty, while on ARC, GPQA and MMLU the gain never exceeds four points, a pattern consistent across all architectures and hidden by average scores. The Proposer‑Critic architecture captures the arithmetic gains, scales most steeply, and at the largest budget outperforms every other design (intervals exclude zero), yet it ranks among the weakest elsewhere, and no single architecture dominates all tasks. To explain these trajectories the authors introduce an exact generate‑transform decomposition. Any workflow can be partitioned into proposal coverage and a downstream transform; consequently, any accuracy change splits exactly into an extensive coverage dividend and an intensive transformation change. This diagnosis shows that arithmetic tasks retain coverage headroom that a critic‑guided transform converts into accuracy, whereas multiple‑choice benchmarks either saturate coverage or fail to convert it, and in open‑ended code generation the recovery effect nearly vanishes so accuracy tracks coverage. Even with equal call budgets, token cost varies by a factor of 2.1, meaning extra calls create candidate opportunities that only some architectures and tasks can exploit. Thus, scaling a team is a task‑ and architecture‑specific bet rather than a universal lever.
Review