RESEARCH PAPER

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation

Homanga Bharadhwaj; Debidatta Dwibedi; Abhinav Gupta; Shubham Tulsiani; Carl Doersch; Ted Xiao; Dhruv Shah; Fei Xia; Dorsa Sadigh; Sean Kirmani

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. An instruction-conditioned generator produces a human-motion video plan, which a separate robot policy follows with live observations; predicted futures guide executable actions during deployment. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

Gen2Act turns a language instruction and an initial scene image into a generated human demonstration, then uses that video to condition a separate closed-loop robot policy. A pretrained VideoPoet supplies the demonstration without robot-specific fine-tuning. Auxiliary point-track prediction teaches policy representations to retain motion cues, while deployment predicts actions directly from video features and recent robot observations. Real robot results support improved generalization relative to the reported baselines, but plausible generation does not ensure correct execution, and long-horizon reliability remains limited.

Abstract

An abstract has not been added yet.

Affiliations

Google DeepMind; The Robotics Institute, Carnegie Mellon University; Computer Science Department, Stanford University