RESEARCH PAPER

Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models

Shaunak A. Mehta; Ananya Hazarika; Haochen Zhang; Fan Yang; Ryo Moriyama; Wenkai Li; Yash Patel; Kanata Suzuki

Classification

View four quadrants
Major category
Related resources
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Not assigned

Category review. A 2026 survey of representations, VLA policies, world models and integration interfaces. It is useful context for interpreting WAM scope and evaluation, rather than a pre-2026 foundation or a specific WAM method. The paper is retained for its adjacent resource/method role; WAM joint-prediction/IDM architecture quadrants do not apply to the cataloged contribution. Reading evidence

AT A GLANCE

Contribution

This survey organizes robot learning around representations that encode the environment, VLA policies that generate actions, and world models that predict consequences. Its useful contribution is a vocabulary for tracing how information and feedback cross those interfaces. Integration can remain modular; the authors argue that its value should be demonstrated through improved behavior under uncertainty, distribution shift and temporal dependencies. A small original CALVIN video-prediction diagnostic illustrates hidden-object and collision failures, but supplies no quantitative proof that a unified architecture resolves them.

Abstract

An abstract has not been added yet.

Affiliations

Fujitsu Research of America; Carnegie Mellon University; Fujitsu Limited