World Action Models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. Existing WAMs, however, struggle with long‑horizon prediction because generating dense video rollouts is highly inefficient. Recent approaches that predict only a single future frame avoid full video generation but neglect how to progress toward the goal.
We introduce ProWAM, a progressive world action model that simultaneously predicts actions and an ordered sequence of sparse visual sub‑goals, providing explicit visual anchors throughout task execution. Sub‑goal prediction can be learned from large‑scale action‑free videos, allowing the video backbone to handle complex visual planning and relieving the action policy.
For efficient action generation, ProWAM performs a single forward pass of the video backbone to cache sparse sub‑goal features, eliminating iterative full‑video generation and requiring only lightweight action denoising during replanning.
Across extensive evaluations, ProWAM demonstrates superior out‑of‑distribution robustness. On simulation benchmarks it achieves 85.8% success on LIBERO‑Plus and 75.7% on randomized RoboTwin, outperforming the strongest baseline by up to +35.9% relative gain. On RoboCasa365 it reaches 48.1% success (18.2% on the challenging Composite‑Unseen split), ranking fourth overall. Crucially, in zero‑shot real‑world experiments ProWAM attains 70.0% success, surpassing the strongest baseline’s 55.0% by +15.0% absolute (+27.3% relative).
These results highlight the value of progress‑indexed visual foresight for closed‑loop control. The code and demo are available at https://sii-ferenas.github.io/ProWAM-page.
Review: ProWAM’s use of sparse visual sub‑goals offers a scalable way to guide long‑horizon planning, delivering both efficiency and robustness that bring vision‑driven control closer to real‑world deployment.