Reasoning segmentation aims to interpret implicit textual queries and achieve fine‑grained visual perception, which is essential for human‑computer interaction and embodied agents. Existing pipelines usually let multimodal large language models (MLLMs) first produce an explicit chain‑of‑thought (CoT) and then locate the target. Although intuitive, the extra textual tokens cause attention interference, weakening perception‑token generation and enlarging the effective distance between visual tokens.
To eliminate this interference we introduce LIRSeg, which replaces explicit CoT with a compact set of learnable latent tokens for reasoning segmentation. Training proceeds in two stages: spatial alignment grounds the latent tokens in object‑relevant visual evidence, and GRPO further refines them using segmentation rewards.
From an information perspective we add three complementary mechanisms to make the few latent tokens more informative: extreme‑advantage sampling selects the most informative training signals; decoupled exploration‑stability updates learn complementary representations; latent diversity amplification prevents representational collapse.
Extensive benchmarks show that LIRSeg consistently improves segmentation accuracy while reducing reasoning overhead. Compared with the VisionReasoner baseline, LIRSeg gains absolute gIoU improvements of 4.9% on ReasonSeg, 7.1% on MUSE, and 4.7% on MMR, and cuts reasoning tokens by roughly 16×. Code is released in the supplementary materials.
Review