RESEARCH PAPER

Learning Universal Policies via Text-Guided Video Generation

Yilun Du; Mengjiao Yang; Bo Dai; Hanjun Dai; Ofir Nachum; Joshua B. Tenenbaum; Dale Schuurmans; Pieter Abbeel

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. A text-guided video planner predicts future observations, then an independent inverse model converts the generated plan into executable simulated robot controls. Video planning precedes control through an independent inverse model. Reading evidence

AT A GLANCE

Contribution

UniPi turns a language instruction and current image into a video plan, refines its timing, then uses a separately trained inverse model to produce robot controls. Simulated manipulation results support compositional and multitask transfer. Internet pretraining improves generated real-scene plans, but its reported success is a classifier judgment on imagined final frames, not measured physical execution. [e02, e04, e06, e10, e14, e16]

Abstract

An abstract has not been added yet.

Affiliations

MIT; Google DeepMind; UC Berkeley; Georgia Tech; University of Alberta