RESEARCH PAPER

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

Siyuan Huang; Liliang Chen; Pengfei Zhou; Shengcong Chen; Yue Liao; Zhengkai Jiang; Yue Hu; Peng Gao; Hongsheng Li; Maoqing Yao; Guanghui Ren

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Instruction-conditioned multiview future-video learning supplies predictive backbone features to a distinct action diffusion head in EnerVerse-A. This is a complete predictive policy, while its 4D reconstruction loop is offline data refinement. Reading evidence

AT A GLANCE

Contribution

EnerVerse learns instruction-conditioned, multi-view future-video representations, then conditions a separate action diffusion head on their backbone features. Sparse history supports long tasks; an offline 4D Gaussian Splatting loop refines training videos. Its strongest benchmark average uses depth-rendered auxiliary views, while physical block placement exposes an instruction-following weakness.

Abstract

An abstract has not been added yet.

Affiliations

SJTU; AgiBot; Shanghai AI Lab; CUHK MMLab; LV-NUS Lab