RESEARCH PAPER

What Matters for Latent Actions in Robot Learning

Xizhou Bu; Qingda Hu; Lei Zhou; Lingfeng Zhang; Yingbo Tang; Zihao Liu; Xinyi Tao; Zhiqiang Ma; Qingqiu Huang; Chufeng Tang; Hongbo Wang; Jing Zhang; Jiayi Ma; Hangjun Ye; Wei Li; Xiaoshuai Hao

Classification

View four quadrants
Major category
VLA
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. An inverse/forward model learns video-derived latent-action supervision before language-conditioned robot policy training; deployment predicts physical actions without world-model rollouts. Joint code/action prediction is not joint future/action prediction. Reading evidence

AT A GLANCE

Contribution

This empirical study asks which video-derived latent actions improve robot policies. It compares modeling paradigms, regularization, and action integration under a shared training framework. Raw-frame LAPO remains competitive, but preferred dimensionality and action heads depend on the evaluation. Its strongest deployment evidence is improved Franka manipulation after latent-action tuning of a VLM backbone; the forward predictor supplies training supervision rather than an inference-time planner.

Abstract

An abstract has not been added yet.

Affiliations

Fudan University; Tsinghua University; Shenzhen University of Advanced Technology; Sichuan University; Suzhou Evans Intelligent Technology Co., Ltd.; Morphi Intelligence Technology Co., Ltd.; Wuhan University; Xiaomi EV