NeFut Logo NeFut
Admin Login

[CS.AI] Auditing Evidence Use in Medical LLM Diagnosis

Published at: 2026-07-25 22:00 Last updated: 2026-07-26 07:44
#AI #Machine Learning #Medical

In the medical field, large language models (LLMs) are often evaluated based on their ability to select the correct diagnosis; however, diagnostic accuracy alone does not reflect whether the model has appropriately utilized case evidence.

This paper presents a behavioral audit of evidence use in medical diagnosis. We decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins.

Since medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions may reflect plausible differential diagnoses, while suspicious interactions require robustness checks and clinical review.

We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase datasets. Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, indicating that many evidence interactions are clinically plausible rather than failures.

In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence.

These results suggest that accuracy can obscure candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.

Original Source: https://arxiv.org/abs/2607.20848

[h] Back to Home