RESEARCH PAPER

EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields

Yang, Zhaoyang; Jin, Yurun; Qi, Lizhe; Huang, Cong; Chen, Kai

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. EA-WM is a 2026 neural simulator jointly modeling RGB video and visualized kinematic action fields in two full-depth streams. Numerical action inversion is a diagnostic excluded from its training/inference, so it should not be promoted to an executed WAM policy or generic component. Reading evidence

AT A GLANCE

Contribution

To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics.

Abstract

Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly treat video generation as an auxiliary representation for policy learning. Consequently, they insufficiently explore the inverse problem: leveraging action signals to guide video synthesis, thereby often failing to preserve precise robot spatial geometry and fine-grained robot-object interaction dynamics in the generated rollouts. To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. Rather than injecting joint or end-effector actions as abstract, low-dimensional tokens, EA-WM projects actions and kinematic states directly into the target camera view as Structured Kinematic-to-Visual Action Fields. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics. Evaluated on the comprehensive WorldArena benchmark, EA-WM achieves state-of-the-art performance, outperforming existing baselines by a significant margin.

Affiliations

1 Fudan University; 2 Zhongguancun Academy; 3 Zhongguancun Institute of Artificial Intelligence; 4 University of Science and Technology of China