NeFut Logo NeFut
Admin Login

[CS.AI] Didactic Knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

Published at: 2026-09-23 22:00 Last updated: 2026-09-24 00:40
#AI #Machine Learning #LLM

Medical large language models are usually trained on a mix of didactic data (e.g., textbooks) and clinical data (e.g., electronic health records). We conducted token‑matched experiments that systematically varied the didactic‑to‑clinical ratio and measured how data composition influences performance, capability profiles, and error patterns on knowledge‑intensive and clinic‑oriented tasks.

The results reveal an asymmetric transfer across task types: adding clinical data markedly improves clinic‑oriented tasks while remaining competitive on knowledge‑intensive tasks, whereas increasing didactic data mainly boosts knowledge‑intensive performance. Error analysis indicates a "knowing‑doing" gap, where gains in knowledge recall do not reliably translate into better clinical reasoning.

We also find that modest amounts of clinical data capture most of the gains on EHR‑grounded tasks, and the optimal didactic‑clinical mixture depends on the downstream task’s knowledge and reasoning demands.

These findings suggest that medical LLM data curation should be application‑driven, favoring higher proportions of clinical data for reasoning‑intensive use cases.

Review

Original Source: https://arxiv.org/abs/2609.22161

[h] Back to Home