Large language models (LLMs) are being explored for clinical reasoning, yet it remains unclear whether they can correctly revise judgments as patient evidence changes over time. We evaluated longitudinal belief updating by pairing intensive‑care trajectories extracted from electronic health records.
Across several LLMs, conditioning a prediction on a prior judgment more often increased prediction error than reduced it; the same pattern was replicated for a second clinical endpoint. Controlled interventions revealed two failure modes.
The first mode: with the preceding assessment held constant, models reacted more strongly to worsening respiratory evidence than to matched improvement; this asymmetry persisted after headroom normalization at moderate and strong evidence levels.
The second mode: with current evidence fixed, raising the prior risk from 10 % to 90 % shifted model estimates by about 26.2 percentage points, demonstrating a causal influence of prior beliefs. Prompting did not restore reliable updating.
The Evidence‑Validated Longitudinal Update (EVLU) approach identified fewer but more reliable revisions, exposing a trade‑off between reliability and coverage. These findings establish longitudinal belief updating as a distinct dimension of LLM reliability.
Review