VideoWorld: Exploring Knowledge Learning from Unlabeled Videos
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q4 · Dual-system × IDM
- Architecture
- Dual-system
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDMLatent prediction & JEPA
- Source review status
- Not assigned
Category review. An autoregressive model predicts video and latent dynamics codes, then a separately trained robot IDM converts those predictions into executed actions with new observations closing the loop. A video/dynamics generator supplies predictions to a separately trained robot IDM. Reading evidence
Contribution
VideoWorld predicts compact multi-step dynamics codes alongside video frames, then converts predictions into task operations. On rendered 9×9 Go and simulated robot tasks, this representation substantially improves over video-only prediction. Expert-curated training data, language conditioning and a separately supervised robot action decoder qualify the broader claim of learning solely from unlabeled videos.
Abstract
An abstract has not been added yet.
Affiliations
Beijing Jiaotong University; ByteDance Seed