We introduce SLIM, a pushing benchmark that places several small objects in a scene and provides paired visual and language goals for the same configuration. Evaluating a LeWM latent world model on SLIM reveals a stark contrast: a scripted controller that accesses simulator state solves every tier, while the language‑conditioned model succeeds on less than 1% of trials. Diagnostic probes locate the bottleneck in the encoder – its latent representation is almost action‑insensitive, cannot decode either the pusher or object positions, and rollouts are no better than copying the current latent forward.
To repair this, we add a single inverse‑dynamics auxiliary loss during training. The loss is applied both to encoder latents and to predicted latents via a shared head that is discarded at test time. This restores all probe metrics and lifts success from 0.003 to 0.35 overall (0.16 on the hard pushing tier, where a goal‑agnostic policy scores zero). Control experiments attribute the gain to gradients flowing into the encoder. A cheap action‑sensitivity probe, computable without environment access, serves as an empirical necessary condition: configurations below its threshold consistently fail to plan.
With the repaired latent space, we attach a lightweight language‑goal head that plans directly from sentences without retraining the world model. The head reaches 0.84 success on navigation (visual‑goal oracle 1.00), follows the correct named zone even when swapped with a decoy, and degrades gracefully to unseen nouns. A single goal sentence rarely completes a push, but when the push is expressed as a sequence of stage sentences, the head raises success on medium and hard tiers to 0.25, matching the goal‑frame oracle, and the stage transition can be read from the latent alone.
Review