NeFut Logo NeFut
Admin Login

[CS.AI] Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#algorithm #AI #Machine Learning

Clinical diagnosis agents must decide which test to request next and when to issue a diagnosis or defer. Existing benchmarks usually evaluate accuracy after fixed or unconstrained interactions, leaving stopping reliability implicit. We introduce Cros, a risk‑constrained stopping layer that combines state‑wise error ranking, policy design on disjoint development splits, and LTT‑style exact tests of selective diagnostic error and minimum autonomous coverage for full sequential policies. Its finite‑sample guarantee requires the candidate family, testing rule, and any randomization to be frozen before calibration labels are accessed. On a 1,834‑episode MIMIC‑derived abdominal‑pain benchmark, the full ranker achieves exploratory state‑error AUROC 0.853, compared with 0.715 for maximum class probability and 0.552 for the backbone’s native stop score. On the previously viewed 367‑episode evaluation split, analytically averaging over frozen Cros weights yields 16.9% selective error at 78.8% coverage, cost 5.57, and 0.68 tests, versus 30.8% error at 100% coverage, cost 8.14, and 1.53 tests under native stopping. Forced continuation is non‑monotone: error is 28.3% with HPI alone and 34.3% after full workup. Although the uniform‑weight mixture ablation is cheaper on this split despite missing the locked development margins, Cros nominally satisfies the joint criterion in only 6 of 20 development resplits. Because evaluation labels were inspected during earlier development, these findings provide exploratory feasibility and audit evidence, not a confirmatory safety certificate.

Review

Original Source: https://arxiv.org/abs/2609.09678

[h] Back to Home