NeFut Logo NeFut
中 Admin Login

[CS.AI] COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#optimization #LLM #GPT

No single large language model (LLM) can be uniformly reliable across all queries, which drives the development of multi‑model inference systems. Traditional routing selects an initial model and stops, whereas dense collaboration invokes every candidate model for each query. We observe that collaboration is non‑monotonic: peers can recover failures that no single model solves, but they can also corrupt answers that were initially correct.

To address this tension, we introduce COMED (Controlled Model Escalation for Multi‑LLM Deliberation), a post‑anchor controller that enables selective cross‑model collaboration. COMED makes three key decisions:

We formalize the trade‑off with a rescue‑harm decomposition: overall performance improves when the number of errors rescued by collaboration exceeds the harms introduced by it. Experiments span medical, scientific, and general reasoning benchmarks across 16 open‑weight settings. COMED consistently outperforms fixed and routed anchors, achieving up to +10.7 percentage points on MedQA while invoking fewer models and decoding fewer tokens than dense collaboration.

On the HLE benchmark with frontier models (e.g., GPT‑5.5), COMED raises accuracy from 23.1% to 28.1%, surpassing dense collaboration and setting the new state‑of‑the‑art.

Review

Original Source: https://arxiv.org/abs/2609.26913

[h] Back to Home