RESEARCH PAPER

FLARE: Robot Learning with Implicit World Modeling

Ruijie Zheng; Jing Wang; Scott Reed; Johan Bjorck; Yu Fang; Fengyuan Hu; Joel Jang; Kaushil Kundalia; Zongyu Lin; Loic Magne; Avnish Narayan; You Liang Tan; Guanzhi Wang; Qi Wang; Jiannan Xiang; Yinzhen Xu; Seonghyeon Ye; Jan Kautz; Furong Huang; Yuke Zhu; Linxi Fan

Classification

View four quadrants
Major category
WAMs
Quadrant
Not assigned
Architecture
Not assigned
Prediction paradigm
Not assigned
Source review status
Not assigned

Category review. Future-token alignment and action denoising share the policy DiT, explicitly coupling predicted future representations with action learning. It is an implicit world-action policy rather than an independent generic encoder. Reading evidence

AT A GLANCE

Contribution

FLARE trains a flow-matching robot policy to match future observation embeddings inside its action-denoising transformer. Compact, action-trained visual-language targets supply an auxiliary learning signal, including from action-free human videos. Simulation and real-robot improvements support this training recipe; they do not establish an explicit planner or calibrated world simulator.

Abstract

An abstract has not been added yet.

Affiliations

NVIDIA; University of Maryland, College Park; Nanyang Technological University; University of Texas, Austin