No single large language model (LLM) can be uniformly reliable across all queries, which drives the development of multi‑model inference systems. Traditional routing selects an initial model and stops, whereas dense collaboration invokes every candidate model for each query. We observe that collaboration is non‑monotonic: peers can recover failures that no single model solves, but they can also corrupt answers that were initially correct.
To address this tension, we introduce COMED (Controlled Model Escalation for Multi‑LLM Deliberation), a post‑anchor controller that enables selective cross‑model collaboration. COMED makes three key decisions:
- Anchor self‑consistency: if repeated sampling from a model yields consistent outputs, the answer is accepted directly.
- Router margin: the confidence gap from the router indicates whether the answer is ambiguous.
- Lightweight peer probe: for ambiguous cases, a brief probe is sent to a few peer models to estimate the potential benefit of collaboration.
We formalize the trade‑off with a rescue‑harm decomposition: overall performance improves when the number of errors rescued by collaboration exceeds the harms introduced by it. Experiments span medical, scientific, and general reasoning benchmarks across 16 open‑weight settings. COMED consistently outperforms fixed and routed anchors, achieving up to +10.7 percentage points on MedQA while invoking fewer models and decoding fewer tokens than dense collaboration.
On the HLE benchmark with frontier models (e.g., GPT‑5.5), COMED raises accuracy from 23.1% to 28.1%, surpassing dense collaboration and setting the new state‑of‑the‑art.
Review