World-action driving models improve planning by coupling multimodal reasoning with future prediction, yet their inference cost grows with model size and increasingly conflicts with the real‑time latency demands of vehicle control. Existing acceleration techniques set policies before deployment by reducing token count, layer depth, or sampling steps, leaving residual runtime variation after offline profiling and static scheduling largely untapped.
We observe that the largest admissible compute budget varies systematically with the residual runtime state, and the most recent realized latency provides a direct signal of the available compute slack. Motivated by this, we introduce SlackDrive—a pre‑inference compute allocator that reuses realized latency to choose the compute budget for each control step. SlackDrive profiles latency and planning utility for a small discrete set of budgets once, estimates the online compute state from completed forward passes, and selects the highest‑utility budget predicted to stay within the admissible latency envelope, complementing existing profiling and resource scheduling while preserving the driving backbone and its compute actuator.
Evaluated on NAVSIM v2 with DriveDreamer‑Policy, SlackDrive improves latency‑constrained EPDMS by $21.7\%$ under a stringent latency regime, outperforming the strongest baseline. In contrast, the full‑budget model and pre‑configured token‑pruning baselines exceed the admissible latency envelope under runtime contention.
Review