3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Visual planning & IDM3D multiview modeling
- Source review status
- Not assigned
Category review. A language-conditioned 3D flow predictor generates object-motion futures, then endpoint verification, grasp selection and optimization solve executable robot poses from the predicted flow. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
3DFlowAction learns instruction-conditioned object trajectories from human and robot videos, checks a rendered endpoint with GPT-4o, and converts accepted flow into robot poses through grasp selection and optimization. Its four-task physical evaluation reports 70% success, but transfer depends on rigid grasp geometry and follows task-specific human-video fine-tuning.
Abstract
An abstract has not been added yet.
Affiliations
South China University of Technology; Tencent Robotics X; Hong Kong University of Science and Technology; Pazhou Laboratory