AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q4 · Dual-system × IDM
- Architecture
- Dual-system
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDM
- Source review status
- Not assigned
Category review. A forward model predicts task-conditioned future motion tokens, and a distinct inverse model turns those predictions plus current observations into executable robot actions at each control step. Separate forward motion-token predictor and inverse action model form a prediction-then-control pipeline. Reading evidence
Contribution
AMPLIFY learns a compact vocabulary of visual motion from point tracks, predicts that motion from an image and task instruction, and translates it into robot actions through a separate inverse model. Its strongest evidence concerns scarce target-task action labels, including transfer where target-task videos remain available. Better track prediction and physical task completion are evaluated separately.
Abstract
An abstract has not been added yet.
Affiliations
Georgia Tech; Georgia Tech Research Institute