RESEARCH PAPER

VILP: Imitation Learning with Latent Video Planning

Zhengtong Xu, Qiang Qiu, Yu She

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. A latent video planner generates future observations, and a separate adjacent-frame action mapper extracts controls with receding-horizon feedback. A separate video planner and adjacent-frame action mapper implement sequential video-to-action prediction. Reading evidence

AT A GLANCE

Contribution

VILP learns observation-conditioned future videos in a compressed latent space, decodes them, and maps adjacent frames to actions through a separate low-level policy. This makes short-horizon video planning practical on the tested tasks and lets task videos supply information beyond scarce action labels. Simulation gains depend on data and evaluation protocol; real-robot evidence comprises 15 trials per method.

Abstract

An abstract has not been added yet.

Affiliations

School of Industrial Engineering, Purdue University, West Lafayette, USA; Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, USA