RESEARCH PAPER

Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning

Muyuan Liu; Yue Huang; Zheng Liang; Xiang Gao

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. The action-conditioned JEPA predictor is trained with physical-state and inverse-dynamics auxiliaries, then used by CEM to generate executed goal-reaching actions. The auxiliary IDM is not the deployed action generator. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

SA+IDM trains an action-conditioned JEPA world model with two auxiliary heads: inverse dynamics recovers executed actions, while state alignment predicts measured physical state from consecutive image representations. Deployment uses only the encoder and latent predictor inside CEM planning. State alignment improves all four reported tasks over IDM alone, while the diagnostics challenge average temporal straightening as a sufficient representation-quality criterion (e2–e12).

Abstract

An abstract has not been added yet.

Affiliations

GENISOM AI, Beijing, China