Medical image interpretation is pivotal for diagnosis and care, yet applying general‑purpose multimodal large language models (MLLMs) to this domain usually demands costly domain‑specific fine‑tuning. To address this, we introduce Representation‑Guided In‑Context Learning (RG‑ICL), a training‑free inference framework. RG‑ICL employs frozen encoders to retrieve demonstration cases that are aligned with the query, feeding these examples together with the current input to the model without updating any parameters.
Across eight public datasets covering histopathology, radiology, and retinal fundoscopy, RG‑ICL achieved an average improvement of 20 percentage points in classification and 13 percentage points in visual question answering (VQA) over both no‑context baselines and conventional ICL, often matching or surpassing methods that require additional training.
The study highlights that the relevance of retrieved cases matters more than their quantity: six query‑aligned cases outperformed up to thirty‑two randomly selected ones, while fixed or random selections frequently degraded performance below the baseline. For VQA, aligning reference cases with both image content and question intent yielded further gains.
These findings suggest that, for medical image interpretation, curating the reference cases presented to an MLLM is a practical and efficient alternative to retraining the model, delivering substantial performance boosts with zero additional training cost.
Review