RESEARCH PAPER

UniVLA: Learning to Act Anywhere with Task-centric Latent Actions

Qingwen Bu; Yanting Yang; Jisong Cai; Shenyuan Gao; Guanghui Ren; Maoqing Yao; Ping Luo; Hongyang Li

Classification

View four quadrants
Major category
VLA
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. A separately learned inverse/forward latent-action model labels video for autoregressive VLM pretraining; the adapted policy decodes predicted codes into controls, with world-model planning explicitly left to future work. Reading evidence

AT A GLANCE

Contribution

UniVLA learns a discrete action vocabulary from language-annotated videos, teaches a vision-language policy to predict that vocabulary, and adapts a visual-conditioned head to physical controls. Its distinguishing idea is to separate task-relevant motion from distracting changes before policy pretraining. Strong manipulation and navigation results support transfer, but deployment still needs action-labeled adaptation; future-observation prediction trains the latent representation rather than serving as the evaluated inference-time planner (e-lam, e-policy, e-decode, e-libero, e-nav, e-limitations).

Abstract

An abstract has not been added yet.

Affiliations

The University of Hong Kong; OpenDriveLab; AgiBot