We built a scaffolding‑oriented multi‑agent large language model (LLM) AI Standardized Patient (AI‑SP) training platform. The system comprises a patient agent that conducts simulated dialogue, a tutor agent that delivers Socratic prompts without revealing diagnostic information, and a turn‑level evaluator agent that monitors clinical progress while keeping summative scores hidden.
In a randomized controlled trial we enrolled 100 medical students and randomly assigned them to either a multi‑agent scaffolding condition or a control condition. All participants completed two learning sessions under their assigned condition and then took an examination in a patient‑only environment. Performance was measured with an Objective Structured Clinical Examination (OSCE) rubric.
The study found no significant difference in final diagnostic accuracy between groups, but the multi‑agent AI‑SP system yielded higher overall exam scores, with the most consistent gains in communication, expression of empathy, and specific history‑taking behaviors. These results suggest that dedicated LLM agents improve the process quality of simulated clinical interviews without artificially inflating exam outcomes.
To support future work we release a multi‑expert annotated dataset containing dialogue transcripts, checklist annotations, turn‑level evaluations, and OSCE‑aligned scoring outcomes, intended to facilitate pedagogically grounded AI‑SP system development and AI‑supported clinical reasoning training.
Review: The platform illustrates the promise of coordinated multi‑agent architectures in medical education, offering a scalable pathway for enhancing clinical interview training.