VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
Classification
View four quadrants- Major category
- WAMs
- Architecture
- Dual-system
- Prediction paradigm
- Joint prediction
- Subcategories
- Joint video-action modeling
- Source review status
- Not assigned
Category review. Distinct video and action denoisers synchronize generation, with predicted video latents directly conditioning executable action trajectories. The data-synthesis application still uses a complete video-action generator, not only a dataset or generic backbone. Separate video DiT and action U-Net are synchronized with directional video-to-action conditioning. Reading evidence
Contribution
VAG synthesizes robot videos and action sequences from an initial image and instruction. A video diffusion transformer supplies pooled clean-latent predictions to an action U-Net during synchronized denoising. The evidence supports improved action prediction and simulation replay, plus a small real-robot policy-pretraining benefit. It does not establish guaranteed alignment or general closed-loop control.
Abstract
An abstract has not been added yet.
Affiliations
GigaAI; Zhejiang University; Peking University; Institute of Automation, Chinese Academy of Sciences; Robotics Department, Mohamed bin Zayed University of Artificial Intelligence