Before a product or policy change is shipped, the crucial question is how users will react. Augur rehearses this reaction offline: it first builds a typed knowledge graph from the change documents, then populates a grounded persona market, runs an interaction simulation, and finally returns an auditable decision memo recommending one of five possible actions. We assembled the Gold-50 dataset, consisting of fifty real product and policy episodes whose outcomes are known from the public record, and used it to score the five‑way release verdict. The central finding is methodological and negative: most of the measured gap between frontier cloud models and the open‑weight models we fine‑tune and serve offline is due to an under‑specified evaluation rather than a capability difference. We demonstrate this in three ways. First, the prompt envelope alone can dominate the score: keeping weights, cases, and scorer fixed, a LoRA‑SFT adapter on Qwen3‑32B swings from 0% to 73%. Second, in a matched 2×2 ablation, defining the decision taxonomy in the prompt—without changing the model—lifts every frontier model by +24 to +34 percentage points; under the under‑specified prompt, the offline‑served Qwen3‑32B LoRA‑SFT beats all three frontier models (paired McNemar test, Holm‑corrected), and once the prompt is fair no significant difference is detected. Third, agreement with the distillation teacher rises while accuracy does not, and the full pipeline amplifies a systematic “over‑doom” bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67‑90% of the concerns actually raised by the public, and a pre‑registered ablation locates its value—largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is publicly available.
Review: This work highlights how evaluation design can dominate perceived model performance, urging the community to standardize prompts and task definitions before drawing conclusions about the superiority of frontier models.