RESEARCH PAPER

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

Ruicheng Zhang; Mingyang Zhang; Jun Zhou; Xiaofan Liu; Zunnan Xu; Zhizhou Zhong; Puxin Yan; Haocheng Luo; Xiu Li

Classification

View four quadrants
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. MIND-V decomposes instructions into object/arm trajectories and generates candidate manipulation videos with a specialized CogVideoX-based renderer. Its iterative feedback judges generated futures; the separate robot-policy experiment uses synthetic visual goals. The main deliverable is task-conditioned neural world simulation rather than a generic pretrained backbone. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX