NeFut Logo NeFut
中 Admin Login

[CS.AI] Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

Published at: 2026-10-01 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #LLM

As large language models (LLMs) become increasingly integrated into clinical workflows, assessing their reliability under realistic variations of clinical text is essential. This work investigates clinical triage by comparing LLM recommendations with those of practicing physicians while applying text perturbations that preserve the underlying medical scenario. We introduce a benchmark comprising over 6,000 triage scenarios, 7,000 physician annotations, and 225,000 model responses. Two key findings emerge: first, LLMs are more prone than physicians to suggest unnecessary care in the baseline condition, and this propensity grows when inputs are perturbed; second, LLM recommendations are more sensitive to gender and tone perturbations than human recommendations, which remain comparatively stable. These results demonstrate that LLMs can change their advice in response to clinically irrelevant textual changes, highlighting the need for deployment‑oriented evaluations grounded in expert physician behavior. Review

Original Source: https://arxiv.org/abs/2609.38600

[h] Back to Home