Large language models (LLMs) have shown strong capabilities across many domains, and medicine is a promising target. However, their deployment is hampered by a shortage of visual question‑answering (VQA) datasets that reflect clinical reasoning and precise image‑text alignment. We tapped de‑identified medical images and expert commentaries shared on clinician‑oriented social media, and built a rigorous pipeline that combines an advanced LLM with clinician‑in‑the‑loop verification. This pipeline produced ThoughtMed-1M, a long‑form medical VQA dataset with over one million pairs, designed to capture structured clinical logic and tight image‑text alignment. Training a foundational model on this data, FOLTMed (FOundational LLM Trained on ThoughtMed-1M), achieved state‑of‑the‑art results on 42 medical VQA benchmarks, reaching a macro accuracy of 85.4% and generating more clinically coherent responses on the ThoughtMed-1M test set. Compared with the previous best models, FOLTMed improved factuality and similarity scores by 3%–5%, highlighting a scalable paradigm for clinically grounded multimodal LLM research.
Review