Model-based reinforcement learning (MBRL) attains high sample efficiency by planning within learned latent dynamics, yet its performance drops sharply when faced with unseen visual distractions such as background changes, lighting shifts, or camera moves. Visual disturbances first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts during planning, creating a two‑level vulnerability. VIGOR addresses this by enforcing latent‑space consistency, enabling zero‑shot visual generalization while preserving MBRL’s sample efficiency. VIGOR consists of three interlocking components: (1) asymmetric weak‑to‑strong augmentation, which creates both weak‑only and weak‑to‑strong latent views within a single batch; (2) dynamics‑level consistency, which imposes augmentation‑invariant transition predictions via direct latent regression; (3) encoder‑level stabilization, which prevents encoder drift under the cross‑augmentation supervision from dynamics‑level consistency. Evaluations on the DeepMind Control Suite and Robosuite show VIGOR surpasses state‑of‑the‑art model‑free and model‑based baselines, improving by 3.4% on DMC and 43.6% on Robosuite. Ablations reveal VIGOR’s robustness is augmentation‑agnostic: swapping the default augmentation with alternatives from different perturbation families retains strong generalization, confirming that latent‑space consistency, not the specific augmentation, drives robustness.
Review: By constraining latent dynamics to be consistent across visual augmentations, VIGOR mitigates error accumulation caused by visual noise, achieving zero‑shot adaptation to novel distractions and highlighting a practical route for model‑based RL in real‑world visual settings.