VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
Classification
View four quadrants- Major category
- WAMs
- Architecture
- One Model
- Prediction paradigm
- Joint prediction
- Subcategories
- Joint video-action modeling
- Source review status
- Not assigned
Category review. One transformer jointly denoises future-video latents and executable action chunks with reciprocal attention, and the robot replans from fresh observations. One transformer jointly denoises video and actions with reciprocal attention. Reading evidence
Contribution
VideoVLA adapts CogVideoX-5B into a robot policy that jointly denoises future video latents and executable action chunks, conditioned on language and the current image. Its clearest gains concern novel objects and skills transferred between embodiments. Video supervision matters strongly in the reported ablations, but imagined success exceeds executed success and deployment remains slow.
Abstract
An abstract has not been added yet.
Affiliations
IAIR, Xi’an Jiaotong University; Microsoft Research Asia; Fudan University