RESEARCH PAPER

VideoVLA: Video Generators Can Be Generalizable Robot Manipulators

Yichao Shen; Fangyun Wei; Zhiying Du; Yaobo Liang; Yan Lu; Jiaolong Yang; Nanning Zheng; Baining Guo

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
Joint prediction
Source review status
Not assigned

Category review. One transformer jointly denoises future-video latents and executable action chunks with reciprocal attention, and the robot replans from fresh observations. One transformer jointly denoises video and actions with reciprocal attention. Reading evidence

AT A GLANCE

Contribution

VideoVLA adapts CogVideoX-5B into a robot policy that jointly denoises future video latents and executable action chunks, conditioned on language and the current image. Its clearest gains concern novel objects and skills transferred between embodiments. Video supervision matters strongly in the reported ablations, but imagined success exceeds executed success and deployment remains slow.

Abstract

An abstract has not been added yet.

Affiliations

IAIR, Xi’an Jiaotong University; Microsoft Research Asia; Fudan University