Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Neural world simulators
- Source review status
- Not assigned
Category review. The technical report supplies state-conditioned humanoid visual world predictors for simulation/forecasting, using separate RGB-generation and discrete-token prediction models rather than executable action generation. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
Team Revontuli uses two separate, state-conditioned predictors: a LoRA-adapted Wan video model for future RGB frames and a spatio-temporal Transformer for future discrete tokens. Both lead the reported challenge leaderboard, but their scores measure conditional prediction rather than executed humanoid control.
Abstract
An abstract has not been added yet.
Affiliations
Aalto University; University of Edinburgh; Deep Render; DataCrunch; University of Helsinki