VILP: Imitation Learning with Latent Video Planning
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q4 · Dual-system × IDM
- Architecture
- Dual-system
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDM
- Source review status
- Not assigned
Category review. A latent video planner generates future observations, and a separate adjacent-frame action mapper extracts controls with receding-horizon feedback. A separate video planner and adjacent-frame action mapper implement sequential video-to-action prediction. Reading evidence
Contribution
VILP learns observation-conditioned future videos in a compressed latent space, decodes them, and maps adjacent frames to actions through a separate low-level policy. This makes short-horizon video planning practical on the tested tasks and lets task videos supply information beyond scarce action labels. Simulation gains depend on data and evaluation protocol; real-robot evidence comprises 15 trials per method.
Abstract
An abstract has not been added yet.
Affiliations
School of Industrial Engineering, Purdue University, West Lafayette, USA; Elmore Family School of Electrical and Computer Engineering, Purdue University, West Lafayette, USA