NeFut Logo NeFut
中 Admin Login

[CS.AI] VLALight: Lightweight Vision-Language-Action Models for Emergency-Aware Traffic Signal Control

Published at: 2026-09-29 22:00 Last updated: 2026-09-30 01:41
#AI #Machine Learning #optimization

Traffic signal control (TSC) is essential for alleviating urban congestion. Recent vision‑language models (VLMs) enable richer interpretation of intersection scenes, opening new possibilities for visual‑context‑aware TSC. Existing pipelines suffer from loose coupling and repeated image‑to‑text conversions, which discard fine‑grained visual details and introduce substantial latency. VLALight addresses these issues with a lightweight end‑to‑end vision‑language‑action framework that directly maps multi‑view intersection observations and signal‑phase information to discrete signal actions. It fuses directional camera feeds into a single visual input and uses textual instructions to align each view with traffic movements and signal phases, allowing direct action prediction with a compact 0.5 B‑parameter model, without intermediate descriptions or handcrafted traffic‑state representations. Experiments show VLALight achieves the best emergency‑vehicle service among all baselines, reducing pooled emergency waiting time by 21.1% compared to the cascaded VLMLight, while running in real time on local hardware and generalizing to unseen intersection topologies and traffic‑flow patterns.

Review

Original Source: https://arxiv.org/abs/2609.30709

[h] Back to Home