UniVLA: Learning to Act Anywhere with Task-centric Latent Actions
Classification
View four quadrants- Major category
- VLA
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Latent action pretrainingAutoregressive VLA
- Source review status
- Not assigned
Category review. A separately learned inverse/forward latent-action model labels video for autoregressive VLM pretraining; the adapted policy decodes predicted codes into controls, with world-model planning explicitly left to future work. Reading evidence
Contribution
UniVLA learns a discrete action vocabulary from language-annotated videos, teaches a vision-language policy to predict that vocabulary, and adapts a visual-conditioned head to physical controls. Its distinguishing idea is to separate task-relevant motion from distracting changes before policy pretraining. Strong manipulation and navigation results support transfer, but deployment still needs action-labeled adaptation; future-observation prediction trains the latent representation rather than serving as the evaluated inference-time planner (e-lam, e-policy, e-decode, e-libero, e-nav, e-limitations).
Abstract
An abstract has not been added yet.
Affiliations
The University of Hong Kong; OpenDriveLab; AgiBot