RESEARCH PAPER

Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising

Jun Guo; Qiwei Li; Peiyan Li; Zilong Chen; Nan Sun; Yifei Su; Heyun Wang; Yuan Zhang; Xinghang Li; Huaping Liu

Classification

View four quadrants
Major category
WAMs
Architecture
One Model
Prediction paradigm
Joint prediction
Source review status
Not assigned

Category review. A shared video/state/action DiT jointly predicts robot actions and future RGB-D/state observations, with asynchronous denoising releasing executable actions before video completion. Shared bidirectional video/state/action DiT jointly generates future representations and controls. Reading evidence

AT A GLANCE

Contribution

X-WAM adapts a pretrained video diffusion transformer to predict robot actions, future states and multi-view RGB-D observations together. A one-way depth branch adds geometric supervision, while asynchronous denoising releases actions before completing video generation. The paper reports strong simulated manipulation and reconstruction results plus a small physical earphone-packing evaluation. Its central tradeoff is useful spatial supervision without depth decoding during every action step; limited temporal context and delayed control remain unresolved.

Abstract

An abstract has not been added yet.

Affiliations

Tsinghua University; Xiaomi Robotics; Peking University; CASIA