RESEARCH PAPERYear 2025

Latent Action Pretraining from Videos

Seonghyeon Ye; Joel Jang; Byeongguk Jeon; Sejune Joo; Jianwei Yang; Baolin Peng; Ajay Mandlekar; Reuben Tan; Yu-Wei Chao; Bill Yuchen Lin; Lars Liden; Kimin Lee; Jianfeng Gao; Luke Zettlemoyer; Dieter Fox; Minjoon Seo

Classification

View four quadrants
Major category
Components of WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. LAPA learns an explicit quantized inverse-transition representation that labels actionless video, then uses these latent-action targets to pretrain a separate VLA. The manuscript explicitly identifies this latent action model as an action-representation component. Retention refers to that reusable LAM, not a claim of online world-model planning. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX