RESEARCH PAPER

LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models

Zuolei Li; Xingyu Gao; Xiaofan Wang; Jianlong Fu

Classification

View four quadrants
Major category
VLA
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. An observed-transition teacher learns scene/motion latents with future-image/action reconstruction, then distills them into a current-observation VLA and separate action expert. The deployed policy does not run the teacher as a world planner. Reading evidence

AT A GLANCE

Contribution

LatBot learns scene and motion latents from language-conditioned robot and human manipulation videos, supervises them through future-image and physical-action decoding, then distills them into a VLA student. An action expert converts the student's features into executable commands. The strongest evidence concerns downstream manipulation success; universal physical understanding and preserved reasoning remain broader interpretations of these results (e03–e06, e12–e16).

Abstract

An abstract has not been added yet.

Affiliations

Institute of Microelectronics, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Microsoft Research