RESEARCH PAPER

Generative World Modelling for Humanoids: 1X World Model Challenge Technical Report

Riccardo Mereu; Aidan Scannell; Yuxin Hou; Yi Zhao; Aditya Jitta; Antonio Dominguez; Luigi Acerbi; Amos Storkey; Paul Chang

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Not assigned

Category review. The technical report supplies state-conditioned humanoid visual world predictors for simulation/forecasting, using separate RGB-generation and discrete-token prediction models rather than executable action generation. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence

AT A GLANCE

Contribution

Team Revontuli uses two separate, state-conditioned predictors: a LoRA-adapted Wan video model for future RGB frames and a spatio-temporal Transformer for future discrete tokens. Both lead the reported challenge leaderboard, but their scores measure conditional prediction rather than executed humanoid control.

Abstract

An abstract has not been added yet.

Affiliations

Aalto University; University of Edinburgh; Deep Render; DataCrunch; University of Helsinki