NeFut Logo NeFut
中 Admin Login

[CS.AI] Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

Published at: 2026-10-02 22:00 Last updated: 2026-10-06 12:11
#algorithm #Machine Learning #Artificial Intelligence

Vision‑Language Model (VLM) agents equipped with world modeling show strong capabilities for complex reasoning and long‑horizon planning while reducing reliance on costly real‑world interactions.

Existing approaches mainly rely on prospective simulation to predict the outcomes of candidate actions. This forward‑only paradigm lacks constraints for verifying whether an action is causally consistent with the observed state transition, often leading to plausible‑looking but physically incoherent behaviors.

We introduce Retrospective World Modeling, enabling agents to reason backward by estimating the attribution distribution $P(\hat{a}_t|s_t, s_{t+1})$, which identifies the action most likely to have caused a given transition. This distribution captures the explanatory power of actions for specific state changes.

Based on this capability, we define the Self‑Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy’s action and the retrospective explanation. A higher SCR indicates that the selected action aligns better with the inferred cause of the transition.

Integrating SCR into reinforcement learning provides dense, transition‑level feedback, steering agents toward behaviors that are both task‑effective and physically grounded.

Extensive experiments across diverse agentic tasks demonstrate that our retrospective approach substantially improves policy robustness and generalization compared to prospective‑only world modeling baselines.

Review

Original Source: https://arxiv.org/abs/2609.39101

[h] Back to Home