How to Learn from What a Human Would Avoid? Intervention-Aware World Models with Real-World RL for Dexterous Manipulation
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Outside quadrants
- Architecture
- Pending verification
- Prediction paradigm
- Other mechanisms
- Subcategories
- Policy post-training & WM-RLDexterous manipulation
- Source review status
- Not assigned
Category review. Action-conditioned latent dynamics and takeover/reward predictions shape online residual-policy learning for executed dexterous manipulation; this is model-assisted control with separate modules. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence
Contribution
WHIRL learns when a human would take over a dexterous robot and uses that prediction to steer residual-policy training. A frozen imitation prior supplies nominal behavior; an action-conditioned, one-step latent world model supports critic learning and actor risk shaping. Five real-robot tasks show higher autonomous success, but the evidence is restricted to one training seed, one operator and familiar workspace regions.
Abstract
An abstract has not been added yet.
Affiliations
HKUST (Guangzhou); Italian Institute of Technology; Zhejiang University