AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Source review status
- Not assigned
Category review. Coupled modality streams generate future RGB, spatial value maps and robot actions; predicted map information conditions actions, and subsequent RL refines control. Reading evidence
Contribution
AIM jointly generates future RGB observations, spatial contact maps and continuous robot actions. A masked mixture-of-transformers architecture makes the predicted maps the action head's only route to future information. Supervised learning uses simulator-derived contact labels; post-training freezes world/value prediction and refines actions using map responses plus task rewards. The paper reports 94.0% Easy and 92.1% Hard success on RoboTwin 2.0, but the small post-training improvement, missing mechanism ablations and inconsistent schematic require careful interpretation.
Abstract
An abstract has not been added yet.
Affiliations
INFIFORCE Intelligent Technology Co., Ltd., Hangzhou, China; The University of Hong Kong, Hong Kong SAR, China; Shanghai Jiao Tong University, Shanghai, China