RESEARCH PAPER

Towards a Generalizable Bimanual Foundation Policy via Flow-based Video Prediction

Chenyou Fan; Fangzheng Yan; Chenjia Bai; Jiepeng Wang; Chi Zhang; Zhen Wang; Xuelong Li

Classification

View four quadrants
Major category
WAMs
Architecture
Dual-system
Prediction paradigm
IDM
Source review status
Not assigned

Category review. Instruction-conditioned flow and RGB predictors supply future visual goals to a separately trained goal-conditioned action diffusion controller, with fresh goals generated during execution. Separate flow/video predictors and goal-conditioned action model compose a visual-goal-to-action pipeline. Reading evidence

AT A GLANCE

Contribution

CogRobot adapts video diffusion to bimanual control through three learned components: instruction-conditioned flow prediction, flow-conditioned RGB prediction, and a task-specific goal-reaching action policy. Flow offers an intermediate description of motion without requiring action labels for the video models. The strongest physical result is Pull Box success of 0.75 versus 0.05 for DP, but the evidence does not establish a universal action policy or broad unseen-task transfer (e02-framing, e05-video, e06-controller, e09-real, e13-limit).

Abstract

An abstract has not been added yet.

Affiliations

Institute of Artificial intelligence (TeleAI), China Telecom; Northwestern Polytechnical University; Hong Kong University of Science and Technology