NeFut Logo NeFut
Admin Login

[CS.AI] Characterizing Text Branch Sensitivity in Medical Vision-Language Segmentation via Evidence Decoupling

Published at: 2026-09-11 22:00 Last updated: 2026-09-12 06:35
#AI #Machine Learning #Neural

Pre‑trained vision‑language models have improved medical image segmentation by incorporating clinical text, yet the actual contribution of text to pixel‑level predictions remains unclear.

This work conducts a systematic study of text’s role in multimodal medical segmentation. We first compare several common fusion strategies and observe that segmentation performance is largely insensitive to the choice of fusion module.

To probe modality interaction, we introduce the Evidence Decoupling Decoder (EDD), built on evidential deep learning and deep supervision. EDD maintains competitive segmentation accuracy while decomposing image evidence and text‑modulated evidence throughout the decoder, serving as an internal representation analysis tool.

Experiments on four public datasets show divergent sensitivity. Removing text causes catastrophic drops on BUSI and BTMRI, indicating strong reliance; on ISIC and Kvasir‑SEG the impact is marginal.

Further analysis reveals that text influences predictions mainly via global semantic modulation rather than independent spatial localization, and the semantic components driving sensitivity differ across datasets.

These findings deepen our understanding of modality interaction in multimodal medical segmentation and offer practical guidance for future model design.

Review: The paper delivers a clear experimental protocol and an interpretable decoder, pointing out when and how textual cues should be leveraged in medical vision‑language systems.

Original Source: https://arxiv.org/abs/2609.02663

[h] Back to Home