RESEARCH PAPER

DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge

Wenyao Zhang; Hongsi Liu; Zekun Qi; Yunnan Wang; Xinqiang Yu; Jiazhao Zhang; Runpei Dong; Jiawei He; Fan Lu; He Wang; Zhizheng Zhang; Li Yi; Wenjun Zeng; Xin Jin

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. A shared backbone predicts motion, geometry and semantic future latents that condition action diffusion; explicit visual decoders can be omitted while anticipatory latent conditioning remains. Reading evidence

AT A GLANCE

Contribution

DreamVLA trains a shared backbone to anticipate motion-relevant regions, depth and semantic features, then conditions action diffusion on its latent predictions. Explicit visual decoders disappear at inference. It reports 4.44 completed instructions on CALVIN ABC-D and 76.7% real-robot success under an attempt-limited protocol. Its lesson is selective forecasting, tempered by unresolved mask, configuration and table inconsistencies.

Abstract

An abstract has not been added yet.

Affiliations

SJTU; EIT; THU; Galbot; PKU; UIUC; USTC