RESEARCH PAPER

VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis

Xiaolei Lang; Yang Wang; Yukun Zhou; Chaojun Ni; Kerui Li; Jiagang Zhu; Tianze Liu; Jiajun Lv; Xingxing Zuo; Yun Ye; Guan Huang; Xiaofeng Wang; Zheng Zhu

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
Joint prediction
Source review status
Not assigned

Category review. Distinct video and action denoisers synchronize generation, with predicted video latents directly conditioning executable action trajectories. The data-synthesis application still uses a complete video-action generator, not only a dataset or generic backbone. Separate video DiT and action U-Net are synchronized with directional video-to-action conditioning. Reading evidence

AT A GLANCE

Contribution

VAG synthesizes robot videos and action sequences from an initial image and instruction. A video diffusion transformer supplies pooled clean-latent predictions to an action U-Net during synchronized denoising. The evidence supports improved action prediction and simulation replay, plus a small real-robot policy-pretraining benefit. It does not establish guaranteed alignment or general closed-loop control.

Abstract

An abstract has not been added yet.

Affiliations

GigaAI; Zhejiang University; Peking University; Institute of Automation, Chinese Academy of Sciences; Robotics Department, Mohamed bin Zayed University of Artificial Intelligence