CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion
Classification
View four quadrants- Major category
- WAMs
- Architecture
- Dual-system
- Prediction paradigm
- Joint prediction
- Subcategories
- Joint video-action modeling
- Source review status
- Not assigned
Category review. Two coupled video/action DiTs exchange information through bridge attention and co-generate futures and robot controls, optionally followed by action refinement. Separate video and action DiTs exchange information through bridge attention. Reading evidence
Contribution
CoVAR co-generates instruction-conditioned video and robot actions through separate, interacting diffusion transformers. Bridge Attention connects a pretrained video branch to an action branch; a UNet decodes actions, with an additional refinement model for LIBERO90. Reported simulation and UR5 successes support the complete system, while refinement dependence, incomplete evaluation details, and roughly four-second sequence generation limit broader conclusions.
Abstract
An abstract has not been added yet.
Affiliations
University of Freiburg; Ludwig Maximilian University of Munich; Munich Center for Machine Learning (MCML); Technical University of Munich; Huawei Heisenberg Research Center (Munich)