VANTAGE-Bench is the first benchmark that explicitly measures how Vision-Language Models (VLMs) perform in Infrastructure AI scenarios. Unlike consumer‑centric embodied AI that focuses on subject‑oriented videos, Infrastructure AI relies on fixed cameras to provide open‑loop insights such as safety monitoring and operational logging.
The benchmark spans three operational domains—Logistics, Transportation, and Smart Spaces—and unifies image and video evaluation across four capability pillars: semantic, spatial, temporal, and spatio‑temporal. It expands beyond multiple‑choice to eight task formulations, including dense captioning and spatio‑temporal grounding, and introduces a single‑pass trajectory protocol for Single Object Tracking, the first evaluation of its kind on fixed‑camera infrastructure video, scored against specialist trackers.
Annotations cover 3,346 media assets: 3,342 video‑task annotations, 4,281 image‑grounding annotations, and 27,404 detection boxes. Zero‑shot evaluation of 17 models reveals that the shortfall relative to consumer‑centric benchmarks is concentrated rather than universal. Event verification, referring expressions, and temporal localization drop roughly 9–24 points, while Video Question Answering stays within 5.3 points of VideoMME and 2D spatial pointing shows no gap against BLINK.
The temporal pillar is the weakest: the best temporal localization reaches only 55.7 % mIoU and dense video captioning peaks at 37.3 % SODA_c. In tracking, frontier models are within about 5 points of specialist trackers on short horizons but diverge markedly as the horizon extends. Open‑weight models lead 2D object localization outright, indicating that neither model scale nor proprietary access explains the observed pattern.
Data, evaluation harness, and leaderboard are available at https://vantage-bench.org/
Review