Mechanistic interpretability is increasingly used to steer activations, remove circuits and monitor safety. An internal estimator may be accurate on average yet still pick a poor action. ObserverBench is a benchmark framework that tests whether an internal estimator—an observer—is suitable for the intervention, control or safety task it directs. Each task fixes the model, information boundary, allowed actions, decision rule, held‑out cases and loss. The benchmark reports estimation accuracy separately from the loss caused by the chosen action; both are required. Theory shows that in closed‑loop control observer errors matter at the start point and along directions reachable by the allowed intervention. Experiments on circuit‑intervention tasks in GPT‑2‑small and Qwen2.5‑7B reveal that pairwise observers predict unseen effects more accurately but do not always choose better actions; observers trained on action loss select lower‑loss actions. In safety triage, a score that perfectly separates violations can allocate a fixed intervention budget poorly when violation costs differ. Across Qwen2.5‑7B, Gemma‑2‑9B‑it and prospectively frozen Qwen3.5‑9B APPS tasks, AUROC can rank monitors differently from deployment loss, and the best information source changes across models. Sparse SAE readouts also lag behind their layer‑matched dense controls on the reported Qwen panels, due to disclosed activation‑density or checkpoint mismatches. ObserverBench provides fixed task contracts, runnable baselines and table‑based submissions for evaluating interpretability methods through the actions they enable.
Review