RESEARCH PAPER

Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining

Jie Cheng; Ruixi Qiao; Yingwei Ma; Binhua Li; Gang Xiong; Qinghai Miao; Yongbin Li; Yisheng Lv

Classification

View four quadrants
Major category
Foundational work
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. This 2024 general model-based RL work integrates world and conservative value learning in a shared transformer and uses imagined beam search for Atari control. It is a pre-2026 world-model/RL antecedent rather than a specialized contemporary robot WAM. Reading evidence

AT A GLANCE

Contribution

JOWA learns an Atari world model and distributional Q-function through one shared transformer, then searches short imagined futures to choose actions. Its strongest evidence is improved aggregate game return and data-efficient offline adaptation; neither universal game-wise scaling nor unconditional planning optimality is established (e02, e03, e04, e07, e08, e11, e12).

Abstract

An abstract has not been added yet.

Affiliations

State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Alibaba Group