Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Q4 · Dual-system × IDM
- Architecture
- Dual-system
- Prediction paradigm
- IDM
- Subcategories
- Visual planning & IDM
- Source review status
- Not assigned
Category review. Instruction-conditioned flow and RGB predictors supply future visual goals to a separately trained goal-conditioned action diffusion controller, with fresh goals generated during execution. Separate flow/video predictors and goal-conditioned action model compose a visual-goal-to-action pipeline. Reading evidence
Contribution
CogRobot adapts video diffusion to bimanual control through three learned components: instruction-conditioned flow prediction, flow-conditioned RGB prediction, and a task-specific goal-reaching action policy. Flow offers an intermediate description of motion without requiring action labels for the video models. The strongest physical result is Pull Box success of 0.75 versus 0.05 for DP, but the evidence does not establish a universal action policy or broad unseen-task transfer (e02-framing, e05-video, e06-controller, e09-real, e13-limit).
Abstract
An abstract has not been added yet.
Affiliations
Institute of Artificial intelligence (TeleAI), China Telecom; Northwestern Polytechnical University; Hong Kong University of Science and Technology