NeFut Logo NeFut
中 Admin Login

[CS.AI] Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Choosing when to forward a query from a small language model to a larger one requires a cheap pre‑call signal that predicts whether escalation will help. Semantic entropy, originally proposed for hallucination detection, measures how much sampled answers disagree in meaning; high disagreement often flags an unreliable response. We evaluated it on three benchmarks across two model families.

On GSM8K, pairing a small and a large model about twelve times different in size, semantic entropy consistently separated the small model's errors (AUROC 0.871) and boosted routed accuracy by up to nine points over random escalation at equal cost. An earlier promising result on a synthetic benchmark turned out to be misleading: a simple rule based solely on question difficulty, without any model, reproduced semantic entropy almost exactly. The main contribution of this paper is a checklist that catches such artefacts before they are reported as genuine findings.

Key observations include:

  1. Scoring a cheap, question‑only difficulty estimate alongside any signal reveals whether the signal adds real information or merely tracks apparent difficulty.
  2. Two reasonable definitions of “escalation worked” can yield very different outcomes on the same data.
  3. Some benchmarks leave almost no room for any signal to beat the trivial strategy of always using the large model.
  4. The true cost of live sampling can make routing more expensive than calling the large model directly.
  5. A cheaper alternative reuses cached past outcomes; we show how to predict its performance on a new dataset and correctly forecasted a collapse from AUROC 0.908 to chance level (0.518) in advance.

Based on these findings we propose a general checklist for evaluating escalation signals.

Review

Original Source: https://arxiv.org/abs/2610.07354

[h] Back to Home