JEPA Policy: Diffusion-Free Imitation Learning via Paired Action and Future Representation Prediction
Classification
View four quadrants- Major category
- WAMs
- Architecture
- One Model
- Prediction paradigm
- Joint prediction
- Subcategories
- Joint video-action modelingLatent prediction & JEPA
- Source review status
- Not assigned
Category review. One mutually visible Transformer predicts executable action chunks and future visual representations together; the paired prediction architecture is supported even though deployment uses no candidate search. A shared mutually visible Transformer predicts actions and future-latent tokens together. Reading evidence
Contribution
JEPA Policy trains a shared Transformer to predict an action chunk and the visual representation observed later in the same demonstration. Two deterministic passes refine both outputs. Its strongest evidence is improved simulated control relative to action-only MIP, supported by topology controls; future error also offers a delayed, task-dependent rollout diagnostic.
Abstract
An abstract has not been added yet.
Affiliations
Anyverse Dynamics