RESEARCH PAPER

Dreamitate: Real-World Visuomotor Policy Learning via Video Generation

Junbang Liang; Ruoshi Liu; Ege Ozguroglu; Sruthi Sudhakar; Achal Dave; Pavel Tokmakov; Shuran Song; Carl Vondrick

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. Task-specific stereo video generation supplies future tool poses, which CAD-based tracking and IK convert into physical robot execution. This is a modular video-to-control method rather than a learned IDM or generic component. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

Dreamitate learns task-specific video generation from human tool demonstrations, then extracts tool poses from synthesized stereo videos for robot execution. It outperforms the tested Diffusion Policy baseline on four physical manipulation tasks, while depending on calibrated cameras, known rigid tools, and slow open-loop generation. The central evidence concerns executed tool trajectories under specified object and scene shifts, rather than unrestricted manipulation or a general action-conditioned simulator.

Abstract

An abstract has not been added yet.

Affiliations

Columbia University; Toyota Research Institute; Stanford University