NeFut Logo NeFut
Admin Login

[CS.AI] A Collective Capability Boundary in Frontier Large Language Models for Guideline-Conformant and Case-Specific Oncology Decision-Making

Published at: 2026-09-01 22:00 Last updated: 2026-09-02 01:40
#AI #Machine Learning #LLM

Large language models (LLMs) achieve high scores on medical knowledge exams, yet real‑world oncology is not a pure knowledge test—it involves selecting guideline pathways, escalation judgments, and making commitments under uncertainty. Existing benchmarks mainly assess factual recall and leave open whether frontier LLMs share decision‑path blind spots that cannot be fixed by model ensembles.

We introduced the Oncology Decision Boundary Benchmark (ODBB), comprising 2,005 oncology decision points drawn from NCCN guidelines and colorectal‑cancer cases. Nine frontier LLMs (four closed‑source, five open‑weight families released between June 2025 and April 2026) were evaluated. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types and was independently validated by two oncologists on a stratified sample of 225 items, achieving Cohen’s weighted $\kappa$ of 0.939 and 0.790.

Treating the nine models as a pooled super‑model, $42.1%$ (Wilson 95% CI 40.0–44.3%) of all items—$35.7%$ of the 1,586 NCCN items and $66.4%$ of the 419 colorectal‑cancer cases—were answered correctly by none. Failures clustered at the stage of choosing between guideline pathways before any reasoning within a pathway, indicating a consistent blind spot in clinical meta‑judgment that likely requires architectural intervention rather than more training data.

Two models tuned for decisiveness (GPT‑5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models, without achieving higher scores. In $3%$–$9%$ of items, models stated the correct next clinical step but did not commit to it—failures of decision, not knowledge.

The takeaway is that model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that a single model can serve as the sole basis for a clinical decision. Progress will require architectures that detect when a model reaches its competence boundary and route the decision to a clinician.

Blogger's Review: This work spotlights a systemic meta‑judgment gap in even the most advanced LLMs, underscoring the need for safety‑aware architectures that know when to defer to human expertise rather than merely chasing higher accuracy.

Original Source: https://arxiv.org/abs/2608.28592

[h] Back to Home