NeFut Logo NeFut
中 Admin Login

[CS.AI] TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#algorithm #AI #Machine Learning

Existing roadside traffic datasets focus on perception, forecasting, and visual question answering, yet they do not assess counterfactual video generation, where a selected actor is altered while the future frames must still respect road topology and unrelated traffic. To fill this gap we introduce TrafficImag, the first benchmark dedicated to counterfactual roadside traffic video generation. TrafficImag comprises 9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor‑centered history‑future samples, together with an executable protocol that supports behavior reasoning, intervention‑aware image editing, and conditional video generation. Each intervention is encoded as an actor‑level program specifying the target actor, intended behavior, legal route, interaction order, and temporal constraints, offering a unified evaluation interface across heterogeneous foundation models. The benchmark measures four validity dimensions: initial‑state correctness, route and behavior validity, interaction consistency, and non‑target preservation; an end‑to‑end counterfactual is considered successful only when all four are satisfied. Results show the strongest reasoner reaches a macro F1 of 80.4%, and the full condition interface raises the best generator’s success rate from 23.3% to 55.0%. Oracle studies reveal that conditional video execution remains the primary bottleneck. TrafficImag provides a reproducible platform for evaluating and diagnosing counterfactual traffic video generation beyond perceptual quality.

Review

Original Source: https://arxiv.org/abs/2609.30722

[h] Back to Home