We built the Oncology Decision Boundary Benchmark (ODBB), comprising 2,005 decision points drawn from NCCN guidelines and colorectal‑cancer cases. Nine frontier large language models released between June 2025 and April 2026 (four closed‑source, five open‑weight families) were evaluated. A fully deterministic scorer—requiring no model inference—classified outputs into 14 failure types and was independently validated by two oncologists on a stratified sample of 225 items, yielding Cohen's weighted κ of 0.939 and 0.790. Treating the nine models as a pooled super‑model, 42.1% of all items (Wilson 95% CI 40.0–44.3%)—35.7% of the NCCN items and 66.4% of the colorectal‑cancer cases—were answered correctly by none. Failures clustered at the stage of choosing between guideline pathways rather than reasoning within a chosen pathway, revealing a consistent blind spot in clinical meta‑judgment that likely demands architectural intervention rather than more data. Two models tuned for decisiveness (GPT‑5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models, without achieving higher scores. In 3–9% of items, models stated the correct next clinical step but did not commit to it—errors of decision, not knowledge. The implication is that model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that a single model can serve as the sole basis for a clinical decision. Progress will require architectures that detect when a model reaches its competence boundary and route the decision to a clinician.
Review