NeFut Logo NeFut
中 Admin Login

[CS.AI] What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

Published at: 2026-10-05 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

Large language models allocate the same amount of compute to every generated token, even though some tokens are intrinsically harder than others. To measure the true per‑token compute, we introduce a Mixture‑of‑Agents (MoA) framework: fifteen language models of increasing capacity, drawn from three families, form a panel. Each agent tries to reproduce a reference token sequence token‑by‑token, conditioned on the correct preceding tokens. The inference cost of the smallest agent that succeeds defines the token’s sufficient compute, an upper bound on the FLOPs the token actually requires.

On three core benchmarks, a 0.5 B parameter agent reproduces 92%‑95% of reference tokens. Across Qwen, OLMo, and R1‑distilled panels, the most expensive 10% of tokens account for 64%‑80% of the estimated FLOPs. On all 500 MATH‑500 problems, the MoA‑derived map enables model routing to cut projected latency from 7.59 s to 5.12 s while slightly improving accuracy, outperforming the best confidence‑routing baseline. For drafting, the MoA map reduces draft tokens by 32.6% and lowers projected latency by roughly 20% compared with fixed‑window drafting at comparable accuracy.

These results reveal substantial allocation headroom, motivating the development of controllers that exploit the sufficient‑compute structure for more efficient inference.

Review

Original Source: https://arxiv.org/abs/2610.02491

[h] Back to Home