RESEARCH PAPER

ω-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

Li, Zhe; Zhang, Zhenzhe; Wei, Yangyang; Zhang, Wenjie; Yuan, Xichen; Zhi, Peiyuan; Li, Gen; Guo, Xinying; Gao, Fengjie; Yang, Jianfei; Zhang, Shanghang

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Verified from primary sources
AT A GLANCE

Contribution

We present ωω-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Real-world experiments on 11 household tasks demonstrate that a single ωω-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.

Abstract

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present ωω-0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, ωω-0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, ωω-0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect ωω-HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single ωω-0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.

Affiliations

1MARS Lab, NTU,2PKU, 3BAAI, 4HKUST(GZ)