FLARE: Robot Learning with Implicit World Modeling
Classification
View four quadrants- Major category
- WAMs
- Quadrant
- Not assigned
- Architecture
- Not assigned
- Prediction paradigm
- Not assigned
- Subcategories
- Latent prediction & JEPA
- Source review status
- Not assigned
Category review. Future-token alignment and action denoising share the policy DiT, explicitly coupling predicted future representations with action learning. It is an implicit world-action policy rather than an independent generic encoder. Reading evidence
Contribution
FLARE trains a flow-matching robot policy to match future observation embeddings inside its action-denoising transformer. Compact, action-trained visual-language targets supply an auxiliary learning signal, including from action-free human videos. Simulation and real-robot improvements support this training recipe; they do not establish an explicit planner or calibrated world simulator.
Abstract
An abstract has not been added yet.
Affiliations
NVIDIA; University of Maryland, College Park; Nanyang Technological University; University of Texas, Austin