Evidence‑based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) mainly focus on superficial label matching, overlooking genuine reasoning. To address this, we introduce the LogiMed‑RoB benchmark, grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries.
LogiMed‑RoB adopts a Hierarchical Logical Consistency (HLC) framework that assesses four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness.
Experiments on ten state‑of‑the‑art LLMs reveal a Error Compounding Effect: the best model achieves 98.88% atomic consistency but its end‑to‑end consistency collapses to 45.13%; several open‑weight architectures drop to near‑zero. We also uncover a systematic evidence‑reasoning gap: even when models retrieve high‑quality evidence, they infer incorrect outcomes in 18.63%‑40.05% of cases, while blind‑guess rates reach 48.28%.
These findings demonstrate that high outcome accuracy can mask critical reasoning flaws, underscoring the need for white‑box logical verification before clinical deployment.
Review