Abstract
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. This work uses mechanistic interpretability to study how robustness-relevant perturbations are represented in WAM activation space. By comparing activations across successful and unsuccessful rollouts, we find that some WAM architectures exhibit low-dimensional linear separability for robustness-critical features, while others do not. This motivates the use of contrastive activation directions for training-free WAM steering.
We also show that local linearity in WAM activation dynamics allows for efficient feedback steering via model-based optimal control, leading to the World-Action Linear Quadratic Regulator (WA-LQR), a minimally-invasive reduced-order LQR controller. Mechanistic evaluations predict strong steerability in the Cosmos-Policy and DiT4DiT models but weak steerability in LingBot-VA, consistent with steering intervention results. On Cosmos-Policy and DiT4DiT, WA-LQR generalizes contrastive directions to new tasks and improves robustness to camera, gripper, and visual-noise perturbations over unsteered and prompt steering baselines.
Blogger's Review: This study provides a fresh perspective on enhancing robustness in World Action Models by integrating mechanistic interpretability with optimal control. The introduction of WA-LQR significantly boosts the flexibility and stability of models in practical applications, making it a compelling avenue for future research exploration.