VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
Classification
View four quadrants- Major category
- VLA
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Latent action pretraining
- Source review status
- Not assigned
Category review. Video-dynamics code prediction pretrains an instruction-conditioned autoregressive model, then a separately supervised CALVIN action head maps its features to executable controls; inference-time video planning is not established. Reading evidence
Contribution
VideoWorld 2 learns discrete visual-dynamics codes with a pretrained diffusion appearance prior, then trains an autoregressive transformer to predict those codes. It improves generated long-horizon craft sequences and transfers latent pretraining to a separately action-supervised CALVIN policy. Its evidence supports benchmark-specific transfer; generated craft success, simulator control and complete appearance disentanglement remain distinct claims.
Abstract
An abstract has not been added yet.
Affiliations
ByteDance Seed; Beijing Jiaotong University