RESEARCH PAPERYear 2025

RLVR-World: Training World Models with Reinforcement Learning

Jialong Wu; Shaofeng Yin; Ningya Feng; Mingsheng Long

Classification

View four quadrants
Major category
WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. RLVR-World trains action-conditioned future-state predictors using verifiable next-state rewards, and its WebArena application uses those predictions in MPC to choose executable actions. It is a world-model training/control method rather than a generic component. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX