Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Visual planning & IDM3D multiview modeling
- Source review status
- Not assigned
Category review. Generated video futures become object-level 3D flow references, which model-based optimization or RL converts into executable robot behavior. The control solver is not automatically a learned IDM. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
Dream2Flow converts generated human-interaction videos into 3D object trajectories, then uses domain-specific optimization or reinforcement learning to make a robot realize them. Its central benefit is an object-level interface across embodiments; its reliability still depends on video geometry, tracking, contact assumptions and the downstream controller.
Abstract
An abstract has not been added yet.
Affiliations
Stanford University