RESEARCH PAPER

EnerVerse-AC: Envisioning Embodied Environments with Action Condition

Yuxin Jiang; Shengcong Chen; Siyuan Huang; Liliang Chen; Pengfei Zhou; Yue Liao; Xindong He; Chiming Liu; Hongsheng Li; Maoqing Yao; Guanghui Ren

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Not assigned

Category review. EVAC predicts action-conditioned multiview observations as a learned environment; an external GO-1 policy supplies actions in evaluation rollouts. Its own output is future imagery, not robot actions. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence

AT A GLANCE

Contribution

EVAC turns robot action sequences into future camera observations using a video diffusion model. Spatial pose maps, temporal action differences and camera rays condition the generated environment. A separate policy can interact with that environment for evaluation, or learn from synthetic trajectories. The strongest evidence is a small policy-data augmentation experiment and agreement with real-robot evaluation trends; physical accuracy remains incompletely measured.

Abstract

An abstract has not been added yet.

Affiliations

AgiBot; SJTU; MMLab-CUHK