Ultrasound is a widely used medical imaging modality, and recent large vision‑language models have shown progress in ultrasound image understanding. However, they lack pixel‑level visual evidence aligned with their semantic predictions, leaving fine‑grained grounding capability unclear.
To address this, we introduce UltraG-Bench, a large‑scale multi‑task benchmark built from 40 public ultrasound segmentation datasets covering 13 anatomical categories. The benchmark comprises three progressive tasks: instruction‑guided segmentation, evidence‑grounded visual question answering, and evidence‑grounded report generation, with 331125, 666779, and 138832 annotations respectively.
Comprehensive evaluation of 14 state‑of‑the‑art models reveals a substantial gap between semantic understanding and pixel‑level localization.
We further propose UltraG‑Agent, which combines the semantic reasoning of a VLM with the ultrasound‑specific segmentation capability of UltraSAM3. Experiments demonstrate that UltraG‑Agent markedly improves both semantic prediction and pixel‑level grounding.
The dataset and code are released at https://github.com/zhuqh19/UltraG-Bench.
Review