Driving models are increasingly grounding their reasoning in causal links, spatial layouts, perceptual cues, and predicted futures. While this makes reasoning more faithful to the scene, it leaves a fundamental question unanswered: what should groundedness mean when the model finally outputs an action? Correctly grounded reasoning alone does not guarantee safe or desirable driving outcomes.
GroundAct starts from a simple premise: driving unfolds through physical entities and their interactions. Consequently, entities become the unit of grounding; a lightweight reference token makes each selected entity's continuous state addressable within symbolic reasoning; and only the interactions between the referenced entities and the evolving proposal adjust the plan. This creates an explicit chain from what the reasoning grounds to what the plan does, which we call grounded planning.
To assess its practical value, GroundAct is evaluated in both open‑loop and closed‑loop settings. Results show strong open‑loop planning across normal, out‑of‑distribution, and safety‑critical scenarios, and closed‑loop simulations extend this evidence to real‑time driving.
Review