Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
Classification
View four quadrants- Major category
- WAMs
- Architecture
- One Model
- Prediction paradigm
- Joint prediction
- Subcategories
- Joint video-action modeling3D multiview modeling
- Source review status
- Not assigned
Category review. A shared video/state/action DiT jointly predicts robot actions and future RGB-D/state observations, with asynchronous denoising releasing executable actions before video completion. Shared bidirectional video/state/action DiT jointly generates future representations and controls. Reading evidence
Contribution
X-WAM adapts a pretrained video diffusion transformer to predict robot actions, future states and multi-view RGB-D observations together. A one-way depth branch adds geometric supervision, while asynchronous denoising releases actions before completing video generation. The paper reports strong simulated manipulation and reconstruction results plus a small physical earphone-packing evaluation. Its central tradeoff is useful spatial supervision without depth decoding during every action step; limited temporal context and delayed control remain unresolved.
Abstract
An abstract has not been added yet.
Affiliations
Tsinghua University; Xiaomi Robotics; Peking University; CASIA