EgoGenesis: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Neural world simulators
- Source review status
- Verified from primary sources
Category review. EgoGenesis resimulates supplied manipulation trajectories into egocentric video using action geometry and scene memory. Its generated video-action pairs train a separate LingBot-VA policy, while the generator itself neither infers nor executes actions. It belongs with learned neural simulators, not general-purpose video backbones. Reading evidence
Contribution
We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms.
Abstract
Egocentric video offers rich manipulation experience for embodied AI, yet collecting diverse egocentric data across scenes, objects, motions, and embodiments remains costly. We present \method, an egocentric world-action simulator that synthesizes controllable, high-quality manipulation videos to expand scarce real-world training data. \method{} builds on a pretrained video generation prior and introduces two geometry-aware conditioning mechanisms. Online Anchored Projective Memory (OAPM) preserves a first-frame 3D scene anchor while periodically refreshing a recent state during autoregressive generation. Action-3D Rotary Position Embedding (A3D-RoPE) encodes end-effector motion with camera-aware 3D rotary coordinates, injecting action geometry into skeleton-to-video cross-attention for precise control. Together, these components improve visual fidelity, geometric stability, and action alignment in long egocentric rollouts. Moreover, augmenting 400 real trajectories with 400 \method-generated trajectories improves out-of-distribution real-robot success from 77\% to 84\% on single-arm tasks and from 53\% to 70\% on dual-arm tasks, demonstrating that the synthesized data substantially improve downstream WAM generalization.
Affiliations
1 Shanghai Jiao Tong University; 2 Alibaba Group; 3 Tianji KernalMind Co., Ltd; 4 The Hong Kong University of Science and Technology; 5 Southeast University; 6 Renmin University of China; 7 The University of Tokyo