Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining
Classification
View four quadrants- Major category
- Foundational work
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Classical world models & model-based RL
- Source review status
- Not assigned
Category review. This 2024 general model-based RL work integrates world and conservative value learning in a shared transformer and uses imagined beam search for Atari control. It is a pre-2026 world-model/RL antecedent rather than a specialized contemporary robot WAM. Reading evidence
Contribution
JOWA learns an Atari world model and distributional Q-function through one shared transformer, then searches short imagined futures to choose actions. Its strongest evidence is improved aggregate game return and data-efficient offline adaptation; neither universal game-wise scaling nor unconditional planning optimality is established (e02, e03, e04, e07, e08, e11, e12).
Abstract
An abstract has not been added yet.
Affiliations
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Alibaba Group