NeFut Logo NeFut
中 Admin Login

[CS.AI] DEEPO: Dual-Entropy Enhanced Policy Optimization for Reducing Hallucination in Multimodal LLMs

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #optimization #LLM

Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its impact on hallucination is uneven. We trace the issue to two weak points in the correction chain from reward to parameter update.

At the rollout level, hard queries with high semantic entropy often yield uniformly wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk peaks.

At the optimization level, confident‑but‑wrong tokens become gradient‑invisible: as a categorical policy sharpens, the norm of the expected score gradient vanishes, so the predictions that need correction receive the weakest updates.

We propose Dual‑Entropy Enhanced Policy Optimization (DEEPO), a dual‑stage enhancement combining signal variance regularization with gradient preconditioning.

Both branches outperform GRPO individually; their interaction yields a statistically significant gain on VideoMMMU—the most complex long‑horizon task in our suite (+$4.0$, 95% CI [1.1, 6.9])—and additive improvements elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.

Review: DEEPO addresses RL’s blind spots in hallucination mitigation by injecting expert guidance on high‑entropy inputs and fine‑tuning gradients, offering a practical route to more reliable multimodal models.

Original Source: https://arxiv.org/abs/2609.28570

[h] Back to Home