Multi‑modal planning promises autonomous driving by offering several plausible behaviors in ambiguous and long‑tail scenarios. Existing works mainly focus on increasing trajectory diversity, enhancing trajectory representations, or reshaping the candidate distribution. Yet we observe a pronounced generation‑evaluation asymmetry: despite strong oracle performance, current planners often fail to reliably pick the best candidate, leaving much potential untapped.
To bridge this gap, we introduce iDriveVLA, a framework that simultaneously expands the candidate trajectory space and provides context‑aware evaluation. The core is a unified trajectory evaluator composed of a Safety‑aware Scorer for quality and risk estimation, and a VLM‑guided Modulator that adaptively weights evaluation criteria based on the scene.
Training follows an oracle‑aligned progressive strategy: first, candidate imitation pre‑training; second, candidate space refinement to boost diversity; finally, semantic ranking alignment to ensure the evaluator’s scores reflect true ordering.
On the public NAVSIM v1 leaderboard, iDriveVLA achieves a new state‑of‑the‑art 94.95 PDMS, surpassing the human‑expert reference.
Review