Learning Massively Multitask World Models for Continuous Control
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Policy post-training & WM-RL
- Source review status
- Not assigned
Category review. Language-conditioned latent dynamics, reward/value models and a policy prior feed explicit CEM/MPPI planning that selects executed continuous-control actions. This is a complete model-based multitask agent. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
Newt extends TD-MPC2 into a language-conditioned agent that learns latent dynamics from demonstrations and online interaction across MMBench. Its practical recipe pretrains the world model and policy prior, retains action supervision, and uses the model to plan. It improves over the evaluated multitask baselines, but specialist policies remain stronger and long open-loop execution is uneven. The evidence concerns simulated, predominantly state-based control.
Abstract
An abstract has not been added yet.
Affiliations
University of California San Diego