RESEARCH PAPER

EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE

Yan, Zexuan; Wu, Yuzhou; Ma, Yue; He, Zonghang; Yin, Kaibo; Tu, Xiaobing; Wang, Yinggui; Ren, Jinkui; Zhang, Xiantao; Wang, Shijian; Liu, Jinghong; Zhang, Linfeng

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. EgoGenesis resimulates supplied manipulation trajectories into egocentric video using action geometry and scene memory. Its generated video-action pairs train a separate LingBot-VA policy, while the generator itself neither infers nor executes actions. It belongs with learned neural simulators, not general-purpose video backbones. Reading evidence

AT A GLANCE

Contribution

We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms.

Abstract

Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.

Affiliations

1 Shanghai Jiao Tong University; 2 Alibaba Group; 3 Tianji KernalMind Co., Ltd; 4 The Hong Kong University of Science and Technology; 5 Southeast University; 6 Renmin University of China; 7 The University of Tokyo