NeFut Logo NeFut
中 Admin Login

[CS.AI] Relevant Evidence Decoding for Audio-Visual Hallucination Mitigation

Published at: 2026-10-06 22:00 Last updated: 2026-10-08 01:25
#AI #Machine Learning #LLM

Audio-Visual Large Language Models (AV‑LLMs) remain prone to cross‑modal hallucinations, where an error in one modality contaminates predictions about another. Contrastive decoding can reduce hallucinations in vision‑language models, but directly applying it to AV‑LLMs overlooks a key issue: different questions require different perceptual evidence—audio, video, or their interaction. We observe that even when a model can answer correctly from a single informative modality, joint audio‑visual inference may weaken the prediction. For instance, when asked “which instrument is heard,” the model may correctly predict violin from audio alone, yet its confidence drops once a video showing a guitar is added.

To address this, we propose Relevant Evidence Decoding (RED), a training‑free decoding method. RED quantifies the predictive support of audio and video beyond the question using pointwise mutual information (PMI) and decomposes their joint contribution into audio, video, and residual interaction components. A question‑only inference pass first determines the required evidence type; then the original audio‑visual prediction is augmented with the corresponding PMI contribution, selectively strengthening relevant evidence.

Across three audio‑visual hallucination benchmarks (CMM, AVHBench, SVHalluc) and three AV‑LLMs, RED improves accuracy over standard decoding by up to 7% on CMM, 6.3% on AVHBench, and 3.8% on SVHalluc, while the average time to first token is about 1.5× that of standard decoding.

Review

Original Source: https://arxiv.org/abs/2610.02976

[h] Back to Home