RESEARCH PAPER

HaWMPO: Hallucination-Aware World Model-based Policy Optimization for Generalist Robot Policy

Zengjue Chen; Peidong Liu; Jiawei Li; Qi Wang

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. A frozen action-conditioned video world model and hallucination/reward scoring guide GRPO post-training of a generalist action policy. The method concerns model-based policy optimization rather than a generic video backbone. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

HaWMPO post-trains OpenVLA-OFT inside a frozen video world model. A separate hallucination detector discounts rewards assigned to unreliable imagined action chunks before GRPO updates the policy. Table 1 reports 63.7% average LIBERO success, but gains vary by suite and the penalty-selection description is inconsistent. Physical testing provides preliminary evidence on two G1 tasks.

Abstract

An abstract has not been added yet.

Affiliations

Joy Future Academy, JD; School of Artificial Intelligence, Jilin University