RESEARCH PAPERYear 2025
RLVR-World: Training World Models with Reinforcement Learning
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Source review status
- Verified from primary sources
Category review. RLVR-World trains action-conditioned future-state predictors using verifiable next-state rewards, and its WebArena application uses those predictions in MPC to choose executable actions. It is a world-model training/control method rather than a generic component. Reading evidence
Contribution
A contribution summary has not been added yet.
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.