RESEARCH PAPER

3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

Hongyan Zhi; Peihao Chen; Siyuan Zhou; Yubo Dong; Quanxi Wu; Lei Han; Mingkui Tan

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. A language-conditioned 3D flow predictor generates object-motion futures, then endpoint verification, grasp selection and optimization solve executable robot poses from the predicted flow. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

3DFlowAction learns instruction-conditioned object trajectories from human and robot videos, checks a rendered endpoint with GPT-4o, and converts accepted flow into robot poses through grasp selection and optimization. Its four-task physical evaluation reports 70% success, but transfer depends on rigid grasp geometry and follows task-specific human-video fine-tuning.

Abstract

An abstract has not been added yet.

Affiliations

South China University of Technology; Tencent Robotics X; Hong Kong University of Science and Technology; Pazhou Laboratory