RESEARCH PAPER
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Joint video-action modelingLatent prediction & JEPA
- Source review status
- Not assigned
Category review. Coupled diffusion streams jointly generate action chunks and future visual embeddings, exchanging information through attention during inference. Reading evidence
Contribution
DUST augments a frozen vision-language backbone with jointly denoised actions and future visual embeddings. Separate streams exchange information through attention; independent noise levels and unequal sampling budgets accommodate their different dynamics. Controlled comparisons support improved manipulation, but do not establish general causal world understanding.
Abstract
An abstract has not been added yet.
Affiliations
Kim Jaechul Graduate School of AI, Korea Advanced Institute of Technology, Seoul, Republic of Korea; RLWRLD, Seoul, Republic of Korea