Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
Classification
View four quadrants- Major category
- VLA
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Source review status
- Not assigned
Category review. A pretrained video network extracts spatiotemporal features from observed images, and a separate language-conditioned diffusion policy generates actions. The source does not describe future prediction in the deployed policy path or a learned future-action generator. Reading evidence
Contribution
Video2Act turns observed video into conditioning for a separate robot action policy. Sobel filtering emphasizes spatial structure in video-model features, while temporal Fourier filtering emphasizes motion. A slow video network refreshes these representations for a faster diffusion action head. The strongest evidence is improved executed manipulation on RoboTwin and a small real-robot evaluation; the speed claims require separating action-chunk throughput from fresh-feedback control.
Abstract
An abstract has not been added yet.
Affiliations
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University; AI²Robotics; Hong Kong University of Science and Technology