Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Visual planning & IDMDexterous manipulation
- Source review status
- Not assigned
Category review. Generated video futures are reconstructed into hand-object references and tracked by learned feedback controllers that output robot commands. This is a specialized modular prediction-to-control system; reference tracking should not automatically be labeled IDM. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
GALATEA turns image-and-language-conditioned manipulation videos into 3-D hand–object references, then learns simulation-based controllers that physically track them. Its central interface retains finger motion and object motion together. Multi-skill experts are distilled into one feedback policy; simulated transfer and 27/40 real-world successes on unseen plans support useful, but limited, generalization (e02, e09, e14, e17).
Abstract
An abstract has not been added yet.
Affiliations
University of California, Berkeley; Sharpa Robotics; The University of Hong Kong