EnerVerse-AC: Envisioning Embodied Environments with Action Condition
Classification
View four quadrants- Major category
- Benchmarks & simulators
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Source review status
- Not assigned
Category review. EVAC predicts action-conditioned multiview observations as a learned environment; an external GO-1 policy supplies actions in evaluation rollouts. Its own output is future imagery, not robot actions. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
EVAC turns robot action sequences into future camera observations using a video diffusion model. Spatial pose maps, temporal action differences and camera rays condition the generated environment. A separate policy can interact with that environment for evaluation, or learn from synthetic trajectories. The strongest evidence is a small policy-data augmentation experiment and agreement with real-robot evaluation trends; physical accuracy remains incompletely measured.
Abstract
An abstract has not been added yet.
Affiliations
AgiBot; SJTU; MMLab-CUHK