RESEARCH PAPER

Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model

John Won; Kyungmin Lee; Huiwon Jang; Dongyoung Kim; Jinwoo Shin

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Coupled diffusion streams jointly generate action chunks and future visual embeddings, exchanging information through attention during inference. Reading evidence

AT A GLANCE

Contribution

DUST augments a frozen vision-language backbone with jointly denoised actions and future visual embeddings. Separate streams exchange information through attention; independent noise levels and unequal sampling budgets accommodate their different dynamics. Controlled comparisons support improved manipulation, but do not establish general causal world understanding.

Abstract

An abstract has not been added yet.

Affiliations

Kim Jaechul Graduate School of AI, Korea Advanced Institute of Technology, Seoul, Republic of Korea; RLWRLD, Seoul, Republic of Korea