RESEARCH PAPER

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

Jeremy A. Collins; Loránd Cheng; Kunal Aneja; Albert Wilcox; Benjamin Joffe; Animesh Garg

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. A forward model predicts task-conditioned future motion tokens, and a distinct inverse model turns those predictions plus current observations into executable robot actions at each control step. Separate forward motion-token predictor and inverse action model form a prediction-then-control pipeline. Reading evidence

AT A GLANCE

Contribution

AMPLIFY learns a compact vocabulary of visual motion from point tracks, predicts that motion from an image and task instruction, and translates it into robot actions through a separate inverse model. Its strongest evidence concerns scarce target-task action labels, including transfer where target-task videos remain available. Better track prediction and physical task completion are evaluated separately.

Abstract

An abstract has not been added yet.

Affiliations

Georgia Tech; Georgia Tech Research Institute