RESEARCH PAPER

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

Wu, Xionghao; Yang, Yijun; Zhou, Shiyang; Sun, Haoze; Liu, Jianhui; Yu, Songsong; Zhang, Jiyao; Li, Wenbo; Wang, Bo; Ma, Guoqing; Song, Lin; Liao, Renjie; Zheng, Shenghe; Tang, Wei; Qi, Xiaojuan; Li, Yanwei; Zhang, Yuan; Tian, Zhuotao; Huang, Haoyang; Duan, Nan

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Verified from primary sources
AT A GLANCE

Contribution

We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video.

Abstract

Robotic manipulation faces a fundamental scaling challenge: robust generalization demands broad physical experience, yet action-labeled robot trajectories are expensive to collect and inherently limited in diversity. Egocentric videos offer a far more scalable source of embodied experience, capturing object interactions, contact dynamics, tool use, and long-horizon behaviors across diverse environments. The central challenge is how to convert this abundant but action-free experience into effective robot control. We introduce ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video. ZimaBlue follows a three-stage training curriculum: it first performs causal embodied video pre-training on large-scale human and robot egocentric videos, then grounds the learned visual dynamics in heterogeneous robot trajectories through video-action mid-training with a unified action representation, and finally specializes the model to a target robot for deployment. To make generative WAMs practical for real-time control, ZimaBluefurther adopts an asynchronous Slow-Fast dual-system architecture, where a high-capacity Slow world model provides generalizable spatiotemporal representations and a lightweight Fast branch enables 30 Hz action prediction on NVIDIA RTX 4090. On real-robot zero-shot evaluations, scaling from target-robot data alone to over 120,000 hours of embodied video improves success from 36.1% to 77.8%. ZimaBlue further delivers strong performance across multiple benchmarks, with particularly pronounced gains on unseen tasks.

Affiliations

Joy Future Academy 1