Recent Video‑LLMs have shown strong performance, yet existing evaluations mainly rely on QA or matching against ground‑truth captions, allowing models to succeed with superficial cues or incomplete annotations. To address this, we introduce VidOmni‑Bench, a benchmark that forces models to verify, sentence by sentence, whether each event in a dense caption is actually supported by the video.
VidOmni‑Bench comprises 500 videos spanning five complexity categories and durations ranging from 4 seconds to 90 minutes. After collecting videos along these axes, we generate dense captions using diverse Video‑LLMs and obtain human‑verified sentence‑level labels; sentences containing incorrect events become hard negatives for evaluation.
Our experiments on VidOmni‑Bench reveal three key findings: (i) Video‑LLMs frequently hallucinate events during dense captioning; (ii) as verifiers they struggle to reliably detect plausible yet wrong event descriptions; (iii) model weaknesses vary with video complexity and duration, exposing distinct, model‑specific bottlenecks.
These results highlight that fine‑grained, cross‑duration video understanding remains a major challenge for current Video‑LLMs, calling for simultaneous improvements in generation accuracy and verification robustness.
Review