RESEARCH PAPER

CoVAR: Co-generation of Video and Action for Robotic Manipulation via Multi-Modal Diffusion

Liudi Yang; Yang Bai; George Eskandar; Fengyi Shen; Mohammad Altillawi; Dong Chen; Ziyuan Liu; Abhinav Valada

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
Joint prediction
Source review status
Not assigned

Category review. Two coupled video/action DiTs exchange information through bridge attention and co-generate futures and robot controls, optionally followed by action refinement. Separate video and action DiTs exchange information through bridge attention. Reading evidence

AT A GLANCE

Contribution

CoVAR co-generates instruction-conditioned video and robot actions through separate, interacting diffusion transformers. Bridge Attention connects a pretrained video branch to an action branch; a UNet decodes actions, with an additional refinement model for LIBERO90. Reported simulation and UR5 successes support the complete system, while refinement dependence, incomplete evaluation details, and roughly four-second sequence generation limit broader conclusions.

Abstract

An abstract has not been added yet.

Affiliations

University of Freiburg; Ludwig Maximilian University of Munich; Munich Center for Machine Learning (MCML); Technical University of Munich; Huawei Heisenberg Research Center (Munich)