World models have achieved remarkable progress in action‑conditioned future prediction, yet they still struggle to generate physically plausible interactions. Existing methods constrain generation with external representations of motion, geometry, or semantics, which require auxiliary estimators or manual annotations and thus limit scalability. By revisiting the training objective we uncover a supervision‑allocation mismatch in the globally averaged mean squared error (MSE) denoising loss: static content dominates the optimization signal, leaving the sparse dynamic‑object regions—crucial for interaction generation—under‑supervised. To address this, we propose IMPACT (Interaction‑aware Model training with Prior‑guided Attention Calibration and Targeting). IMPACT leverages cross‑attention linked to manipulated‑object tokens as an internal spatiotemporal prior. It samples candidate regions from this prior, calibrates them with detached local prediction errors to build an interaction map, and uses the map to reweight denoising supervision. The framework requires no external representations and introduces no inference‑time modifications. Extensive experiments on robot‑arm and human‑hand manipulation across diverse control modalities and DiT backbones show that IMPACT consistently outperforms MSE‑trained baselines, improving interaction fidelity, physical plausibility, and visual quality.
Review