RESEARCH PAPER

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Chenhao Zhang; Hanyu Zhao; Hang Cheng; Tengfei Pan; Long Zeng

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. A separate frozen forward world model imagines candidate action outcomes, and an evaluator/scheduler selects informative imagined experience to post-train executable VLA actions. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

WISE post-trains a VLA action head by selectively imagining alternative behaviors at visually identified interaction states. A separate frozen world model predicts bounded futures; a frozen evaluator ranks them; updates supervise only the first action chunk from a real context. The strongest controlled evidence is the scheduling ablation: higher task success with substantially less imagination computation. The method, results, and reproduction boundaries below trace this conclusion to the primary text.

Abstract

An abstract has not been added yet.

Affiliations

Tsinghua University; Beijing Academy of Artificial Intelligence (BAAI)