RESEARCH PAPER

AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps

Liaoyuan Fan; Zetian Xu; Chen Cao; Wenyao Zhang; Mingqi Yuan; Jiayu Chen

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Coupled modality streams generate future RGB, spatial value maps and robot actions; predicted map information conditions actions, and subsequent RL refines control. Reading evidence

AT A GLANCE

Contribution

AIM jointly generates future RGB observations, spatial contact maps and continuous robot actions. A masked mixture-of-transformers architecture makes the predicted maps the action head's only route to future information. Supervised learning uses simulator-derived contact labels; post-training freezes world/value prediction and refines actions using map responses plus task rewards. The paper reports 94.0% Easy and 92.1% Hard success on RoboTwin 2.0, but the small post-training improvement, missing mechanism ablations and inconsistent schematic require careful interpretation.

Abstract

An abstract has not been added yet.

Affiliations

INFIFORCE Intelligent Technology Co., Ltd., Hangzhou, China; The University of Hong Kong, Hong Kong SAR, China; Shanghai Jiao Tong University, Shanghai, China