Dynamic LLM routers claim to cut inference cost by sending each query to the cheapest model that can answer it correctly. We evaluate six commercial routers across 14 settings on a benchmark covering eight task categories. None of them beats a baseline that randomly chooses between two well‑chosen models at the same cost, and some fall behind by more than ten percentage points.
We trace the gap to four patterns that appear frequently in routers: difficulty blindness, length reversal, semantic matching, and roster suboptimality. The first three are exactly what the standard objective—cost‑accuracy Pareto efficiency on realized costs—rewards: it prefers routing moderately hard queries over the hardest ones, shorter queries over longer ones, and routing based on a query’s source rather than its true difficulty.
We also show that the two assumptions that would justify large rosters—model granularity and model specialization—do not hold empirically. Consequently we propose an evaluation methodology that does not reward these patterns and, as a proof of concept, design a simple two‑model router that avoids all four.
Nevertheless, because a well‑chosen roster already leaves little to gain, the improvement over random routing is modest.
Review