LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
Classification
View four quadrants- Major category
- VLA
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Latent action pretraining
- Source review status
- Not assigned
Category review. An observed-transition teacher learns scene/motion latents with future-image/action reconstruction, then distills them into a current-observation VLA and separate action expert. The deployed policy does not run the teacher as a world planner. Reading evidence
Contribution
LatBot learns scene and motion latents from language-conditioned robot and human manipulation videos, supervises them through future-image and physical-action decoding, then distills them into a VLA student. An action expert converts the student's features into executable commands. The strongest evidence concerns downstream manipulation success; universal physical understanding and preserved reasoning remain broader interpretations of these results (e03–e06, e12–e16).
Abstract
An abstract has not been added yet.
Affiliations
Institute of Microelectronics, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Microsoft Research