World models (WMs) simulate environment transition dynamics, enabling agents to plan over the consequences of their actions. In text‑based settings, fine‑tuning a language model (LM) to act as a WM has become the dominant paradigm. Although Retrieval‑Augmented Generation (RAG) excels in non‑parametric tasks, its use for LM‑based world modeling remains under‑explored.
We conduct a systematic evaluation across five diverse environments—embodied, web navigation, and social scenarios—comparing fine‑tuning and RAG‑based approaches for LM‑based world modeling. Results show that fine‑tuned WMs achieve higher rewards in 15 out of 20 settings, generally outperforming RAG. Both paradigms benefit from richer and more diverse exploration data; RAG is more data‑efficient, while fine‑tuning gains disproportionately from scaling the amount of collected experience.
Focusing on RAG‑based WMs, we devise a counterfactual‑intervention procedure to estimate the retrieval stage error rate, revealing that retrievers frequently surface suboptimal transitions from the experience buffer. To address this, we explore various query‑reformulation strategies and demonstrate that a hierarchical retrieval approach outperforms the traditional pipeline.
Finally, we combine these insights into a hybrid world‑modeling system: the core environment dynamics are captured parametrically, while an actively maintained memory store provides retrieval‑based support. The hybrid system consistently outperforms other methods across multiple environments and models, showcasing its robustness and applicability.
Review