RESEARCH PAPER

Vidar: Embodied Video Diffusion Model for Generalist Manipulation

Yao Feng; Hengkai Tan; Xinyi Mao; Chendong Xiang; Guodong Liu; Shuhe Huang; Hang Su; Jun Zhu

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. A task-conditioned video generator produces futures used at inference by a separately trained masked inverse-dynamics model to produce robot commands. Separate video generator and masked IDM are both used to produce executable actions. Reading evidence

AT A GLANCE

Contribution

Vidar adapts an embodied video generator to a target bimanual robot, then converts predicted imagery into controls using a separately trained masked inverse dynamics model. Low-demonstration real-world results and component ablations support this design; action observability, open-loop execution, and substantial prior training constrain its generality.

Abstract

An abstract has not been added yet.

Affiliations

Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University; Shengshu Tech, Beijing, 100084, China