DexWorldModel: Causal Latent World Modeling towards Automated Learning of Embodied Tasks
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q2 · One Model × IDM
- Architecture
- One Model
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDMLatent prediction & JEPA
- Source review status
- Not assigned
Category review. Shared transformer blocks first predict future DINO features and then generate motor chunks conditioned on those predictions, with execution feedback updating memory. Shared core transformer blocks sequentially predict future features then inverse-dynamics-style actions. Reading evidence
Contribution
DexWorldModel introduces CLWM, which predicts future DINOv3 features and then generates actions conditioned on that prediction. Shared transformer blocks, separate persistent and speculative memories, and asynchronous denoising connect world prediction to robot execution. EmbodiChain supplies synthetic adaptation data. Reported manipulation results are strong, but protocol omissions and a contradictory flow-time convention limit reproducibility.
Abstract
An abstract has not been added yet.
Affiliations
DexForce AI