EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- 3D multiview modelingMemory & long-horizon modeling
- Source review status
- Not assigned
Category review. Instruction-conditioned multiview future-video learning supplies predictive backbone features to a distinct action diffusion head in EnerVerse-A. This is a complete predictive policy, while its 4D reconstruction loop is offline data refinement. Reading evidence
Contribution
EnerVerse learns instruction-conditioned, multi-view future-video representations, then conditions a separate action diffusion head on their backbone features. Sparse history supports long tasks; an offline 4D Gaussian Splatting loop refines training videos. Its strongest benchmark average uses depth-rendered auxiliary views, while physical block placement exposes an instruction-following weakness.
Abstract
An abstract has not been added yet.
Affiliations
SJTU; AgiBot; Shanghai AI Lab; CUHK MMLab; LV-NUS Lab