Dreamitate: Real-World Visuomotor Policy Learning via Video Generation
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Visual planning & IDM
- Source review status
- Not assigned
Category review. Task-specific stereo video generation supplies future tool poses, which CAD-based tracking and IK convert into physical robot execution. This is a modular video-to-control method rather than a learned IDM or generic component. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
Dreamitate learns task-specific video generation from human tool demonstrations, then extracts tool poses from synthesized stereo videos for robot execution. It outperforms the tested Diffusion Policy baseline on four physical manipulation tasks, while depending on calibrated cameras, known rigid tools, and slow open-loop generation. The central evidence concerns executed tool trajectories under specified object and scene shifts, rather than unrestricted manipulation or a general action-conditioned simulator.
Abstract
An abstract has not been added yet.
Affiliations
Columbia University; Toyota Research Institute; Stanford University