RESEARCH PAPER
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Visual planning & IDMLatent prediction & JEPA
- Source review status
- Not assigned
Category review. A shared backbone predicts motion, geometry and semantic future latents that condition action diffusion; explicit visual decoders can be omitted while anticipatory latent conditioning remains. Reading evidence
Contribution
DreamVLA trains a shared backbone to anticipate motion-relevant regions, depth and semantic features, then conditions action diffusion on its latent predictions. Explicit visual decoders disappear at inference. It reports 4.44 completed instructions on CALVIN ABC-D and 76.7% real-robot success under an attempt-limited protocol. Its lesson is selective forecasting, tempered by unresolved mask, configuration and table inconsistencies.
Abstract
An abstract has not been added yet.
Affiliations
SJTU; EIT; THU; Galbot; PKU; UIUC; USTC