Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Latent prediction & JEPA
- Source review status
- Not assigned
Category review. The action-conditioned JEPA predictor is trained with physical-state and inverse-dynamics auxiliaries, then used by CEM to generate executed goal-reaching actions. The auxiliary IDM is not the deployed action generator. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
SA+IDM trains an action-conditioned JEPA world model with two auxiliary heads: inverse dynamics recovers executed actions, while state alignment predicts measured physical state from consecutive image representations. Deployment uses only the encoder and latent predictor inside CEM planning. State alignment improves all four reported tasks over IDM alone, while the diagnostics challenge average temporal straightening as a sufficient representation-quality criterion (e2–e12).
Abstract
An abstract has not been added yet.
Affiliations
GENISOM AI, Beijing, China