Background
As clinical language models evolve, LLM judges increasingly assess whether these models provide overconfident answers under incomplete evidence. However, it's unresolved whether a measured "safety gain" reflects actual behavioral change or the calibration of the judge.
Methods
Using a structured evidence-sufficiency prompt as a test case, we investigated whether it reduces unsafe overconfident answers, how much this effect depends on the scoring judge, and what it costs in terms of helpfulness.
In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper.
The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses incorporated a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded review by three clinicians.
Results
Unsafe overconfidence decreased from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p < 0.001).
Blogger's Review: This study delves into the trade-off between safety and utility in clinical LLMs, revealing the subjective influence of judges on outcomes and providing an essential reference framework for future model evaluations.