RESEARCH PAPER

Learning Counterfactual World Models for Embodied Reasoning under Partial Observability

Todd Y. Zhou; Daniel Zhang

Classification

View four quadrants
Major category
WAMs
Architecture
Pending verification
Prediction paradigm
Other mechanisms
Source review status
Not assigned

Category review. A recurrent action-conditioned latent world model supports MPC search over imagined intervention futures and receding-horizon action execution; contrastive outcome training improves this specific planning system. The described action mechanism is external planning, model-assisted policy optimization, geometric tracking, or video-conditioned control; the source does not establish joint future/action generation or an IDM action decoder. Reading evidence

AT A GLANCE

Contribution

Counterfactual Latent World Models (CLWM) train action-conditioned imagined futures to preserve differences in intervention outcomes even when observations look alike. A recurrent world model supplies latent rollouts to model-predictive control; privileged outcome labels supervise an additional contrastive objective during training. Reported simulation gains reach 11.6 percentage points in navigation success. The evidence supports targeted representation training, while supervision-matched controls and audits of pretrained encoders remain open.

Abstract

An abstract has not been added yet.

Affiliations

Harvard University