NeFut Logo NeFut
中 Admin Login

[CS.AI] MVVBench: Benchmarking 4D Reasoning in Vision-Language Models

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #Neural

Multi‑view video understanding demands the integration of spatial and temporal evidence across several often non‑overlapping camera streams: tracking entities as they move between viewpoints, aligning events over time, and reasoning about latent 4D continuity rather than relying on any single visible frame. We introduce MVVBench, a benchmark for multi‑view video reasoning built from real‑world multi‑camera datasets. Each question is ambiguous both in view and temporal dimensions—no single view in the provided set can answer it, and most questions are also unanswerable from any single moment. Only by jointly reasoning across views and across time does a question become uniquely solvable. MVVBench spans diverse dynamic scenes and probes six capabilities: implicit/explicit attribute identification, implicit/explicit relative distance, relative camera pose, and compositional counting, all with human‑authored QA and rigorous verification. Beyond the benchmark, we conduct an extensive analysis of when and why current vision‑language models succeed or fail, pinpointing errors due to temporal mis‑localization, cross‑view identity breaks, and brittle multi‑hop reasoning. We then study inference‑time elicitation strategies—task‑specific chain‑of‑thought scaffolds and structured cross‑view evidence aggregation—that unlock latent multi‑view competence without retraining, yielding substantial gains. Preliminary evidence shows that reinforcement learning with verifiable rewards can also elicit hidden multi‑view abilities, suggesting training‑time approaches as a promising future direction. In sum, MVVBench offers a rigorous evaluation of 4D multi‑view reasoning and a foundation for progress toward reliable embodied perception.

Review: MVVBench’s design of view‑and‑time ambiguous queries highlights a critical gap in current vision‑language models’ cross‑modal, spatiotemporal reasoning, while providing a clear evaluation framework and practical pathways for improvement, thereby charting a concrete roadmap for future research.

Original Source: https://arxiv.org/abs/2609.30952

[h] Back to Home