NeFut Logo NeFut
中 Admin Login

[CS.AI] Towards VLA-Dreamer: Refining VLA Behavior Using World Models

Published at: 2026-09-30 22:00 Last updated: 2026-10-06 12:11
#AI #Machine Learning #Neural

Vision-Language-Action (VLA) models show strong potential for robot control, yet they require massive amounts of high‑quality imitation data. Moreover, existing VLAs lack an explicit world model, raising doubts about their ability to reason about real‑world dynamics.

We propose a novel architecture that trains a predictive world model directly on the embedding space produced by the VLA’s vision encoder. The hypothesis is that these embeddings retain action‑relevant information and can therefore be used to forecast future states. Unlike conventional pixel‑space world models, the loss is computed in the embedding space, similar to joint‑embedding predictive networks.

After training, the world model can be employed for short‑term planning: given a goal image, the model samples a sequence of VLA actions that transition the current state toward the goal. Experiments will assess the richness of the visual embeddings and test whether the added world model substantially reduces the need for large imitation datasets while enabling on‑the‑fly plan generation during inference.

The overall pipeline consists of: (1) building a predictive model on top of the VLA’s visual encoder outputs; (2) conditioning the embedding predictions on actions to verify whether VLAs implicitly contain a non‑lossy world model; (3) using the trained model for goal‑directed action sampling, thus closing the loop between planning and control.

Review

Original Source: https://arxiv.org/abs/2609.31313

[h] Back to Home