RESEARCH PAPER

VideoWorld 2: Learning Transferable Knowledge from Real-world Videos

Zhongwei Ren; Yunchao Wei; Xiao Yu; Guixun Luo; Yao Zhao; Bingyi Kang; Jiashi Feng; Xiaojie Jin

Classification

View four quadrants
Major category
VLA
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Video-dynamics code prediction pretrains an instruction-conditioned autoregressive model, then a separately supervised CALVIN action head maps its features to executable controls; inference-time video planning is not established. Reading evidence

AT A GLANCE

Contribution

VideoWorld 2 learns discrete visual-dynamics codes with a pretrained diffusion appearance prior, then trains an autoregressive transformer to predict those codes. It improves generated long-horizon craft sequences and transfers latent pretraining to a separately action-supervised CALVIN policy. Its evidence supports benchmark-specific transfer; generated craft success, simulator control and complete appearance disentanglement remain distinct claims.

Abstract

An abstract has not been added yet.

Affiliations

ByteDance Seed; Beijing Jiaotong University