RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer
Classification
View four quadrants- Major category
- Datasets
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Synthetic data & data generation
- Source review status
- Not assigned
Category review. The contribution is offline geometry-preserving multiview augmentation of demonstrated videos, whose existing action labels train a separate ACT policy; it does not generate new actions or plan online. The cataloged contribution is a dataset/data-generation method or evaluation/simulation resource, not the architecture of an evaluated or external policy. Reading evidence
Contribution
RoboTransfer augments robot demonstrations by changing their visual appearance while conditioning video diffusion on the demonstrated geometry. Jointly encoded camera views, metric depth, normals and separate background/object references support multi-view synthesis. A separately trained ACT policy benefits from the augmented observations on two physical manipulation tasks. This is evidence for offline data augmentation, with remaining uncertainty about statistical reliability and physical fidelity.
Abstract
An abstract has not been added yet.
Affiliations
Horizon Robotics; GigaAI; CASIA