RESEARCH PAPER

Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling

Yueru Jia; Jiaming Liu; Shengbang Liu; Rui Zhou; Wanhe Yu; Yuyang Yan; Xiaowei Chi; Yandong Guo; Boxin Shi; Shanghang Zhang

Classification

View four quadrants
Major category
VLA
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. A pretrained video network extracts spatiotemporal features from observed images, and a separate language-conditioned diffusion policy generates actions. The source does not describe future prediction in the deployed policy path or a learned future-action generator. Reading evidence

AT A GLANCE

Contribution

Video2Act turns observed video into conditioning for a separate robot action policy. Sobel filtering emphasizes spatial structure in video-model features, while temporal Fourier filtering emphasizes motion. A slow video network refreshes these representations for a faster diffusion action head. The strongest evidence is improved executed manipulation on RoboTwin and a small real-robot evaluation; the speed claims require separating action-chunk throughput from fresh-feedback control.

Abstract

An abstract has not been added yet.

Affiliations

State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; AI²Robotics; Hong Kong University of Science and Technology