NeFut Logo NeFut
中 Admin Login

[CS.AI] Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #optimization

Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning ability of multimodal large language models, yet existing frameworks rely on coarse sequence‑level rewards and lack fine‑grained supervision of visual grounding. We examine two token‑level metrics: visual dependency (how much a token’s prediction relies on image features) and predictive entropy. Empirical results show: (1) correct reasoning chains exhibit a sharper entropy drop as visual grounding intensifies; (2) pivotal tokens are outliers in the joint visual‑dependency/entropy distribution of correct rollouts, and their misprediction triggers collapse. Motivated by these observations we introduce Token‑level Perception‑grounded Advantage Estimation (TPAE), which measures each token’s statistical consistency with the vision‑entropy patterns of correct rollouts to estimate advantages. The token‑level score modulates the sequence‑level advantage, providing fine‑grained supervision that can be plugged into any RLVR system. Experiments on seven benchmarks demonstrate that TPAE consistently yields more stable and efficient optimization. Code: https://github.com/Zhihan72/TPAE.

Review

Original Source: https://arxiv.org/abs/2609.39168

[h] Back to Home