Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
Classification
View four quadrants- Major category
- Related resources
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Visuomotor policy methods
- Source review status
- Not assigned
Category review. A non-language visuomotor policy uses a human-video prompt, video-model features and retargeted supervision to transfer manipulation skills. It is relevant to embodiment transfer, but the described controller does not explicitly predict future states or perform world-model planning. The paper is retained for its adjacent resource/method role; WAM joint-prediction/IDM architecture quadrants do not apply to the cataloged contribution. Reading evidence
Contribution
A human demonstration video conditions a frozen video model whose features guide a separate diffusion policy for an Xhand robot. Human hand motions are reconstructed and retargeted into the policy's action space, allowing human demonstrations to supervise action learning. The reported transfer concerns objects and skills absent from robot training demonstrations but present in human training data; it does not establish learning an entirely untrained skill from the inference prompt alone.
Abstract
An abstract has not been added yet.
Affiliations
Tsinghua University, China; Shanghai Qi Zhi Institute, China