World models aim to predict what happens next given current environmental conditions, and recent work increasingly equips them with multi‑modal generation, such as visual simulations paired with textual physical states. This capability, however, introduces cross‑modal inconsistency: each output may look plausible in isolation yet disagree—e.g., a video shows a ball rebounding while the text omits the rebound, or both diverge from real physics.
We focus on two explicit failures. Internal misalignment denotes disagreement between a model's generated video and its prediction in another modality (text). External misalignment denotes mismatch between the model's output and an analytically derived physical environment. To make both measurable we define contracts over event, magnitude and timing, and build a physics‑grounded pipeline for comparison in internal and external settings.
We then test whether progressively feeding the model its own contract (the A ladder for internal alignment) or a corrected physical contract (the B ladder for external alignment) narrows the respective gaps. Experiments span four mechanisms and twenty settings, covering varied object interactions and motion patterns.
Across all trials the language component answers every one of the 22 text probes correctly with respect to the true environment, yet the corresponding neutral video often disagrees. This suggests that current unified backbones struggle to achieve correct reasoning, internal consistency and external physical fidelity simultaneously.
Review