RESEARCH PAPER

Grounding Generated Video Plans in Simulation Towards Versatile Dexterous Controllers

Tianyue Wu; Boyuan An; Shuqi Zhao; Heyu Guo; Wanli Xing; Yi Ma; Kaifeng Zhang; Ruihai Wu; Masayoshi Tomizuka

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. Generated video futures are reconstructed into hand-object references and tracked by learned feedback controllers that output robot commands. This is a specialized modular prediction-to-control system; reference tracking should not automatically be labeled IDM. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

GALATEA turns image-and-language-conditioned manipulation videos into 3-D hand–object references, then learns simulation-based controllers that physically track them. Its central interface retains finger motion and object motion together. Multi-skill experts are distilled into one feedback policy; simulated transfer and 27/40 real-world successes on unseen plans support useful, but limited, generalization (e02, e09, e14, e17).

Abstract

An abstract has not been added yet.

Affiliations

University of California, Berkeley; Sharpa Robotics; The University of Hong Kong