RESEARCH PAPER

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos

Zhongwei Ren; Yunchao Wei; Xun Guo; Yao Zhao; Bingyi Kang; Jiashi Feng; Xiaojie Jin

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. An autoregressive model predicts video and latent dynamics codes, then a separately trained robot IDM converts those predictions into executed actions with new observations closing the loop. A video/dynamics generator supplies predictions to a separately trained robot IDM. Reading evidence

AT A GLANCE

Contribution

VideoWorld predicts compact multi-step dynamics codes alongside video frames, then converts predictions into task operations. On rendered 9×9 Go and simulated robot tasks, this representation substantially improves over video-only prediction. Expert-curated training data, language conditioning and a separately supervised robot action decoder qualify the broader claim of learning solely from unlabeled videos.

Abstract

An abstract has not been added yet.

Affiliations

Beijing Jiaotong University; ByteDance Seed