RESEARCH PAPER

FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space

Murray, Michael; Chen, Daphne; Bagaria, Simran; Fortier, Dean; Hellebrekers, Tess; Mullins, Galen; Gajarla, Harshavardhan; Mees, Oier; Cakmak, Maya; Kolobov, Andrey

Classification

View four quadrants
Major category
WAMs
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Verified from primary sources

Category review. FlowDAgger adapts generative policies through a learned noise policy and frozen-base inversion. Its WAM branch inverts the complete joint state/action/value noise tensor and preserves predicted future-state targets while imposing expert action corrections. This is policy/WAM adaptation, not a reusable visual backbone or tokenizer. Reading evidence

AT A GLANCE

Contribution

We present FlowDAgger, a sample- and compute-efficient method for adapting frozen generative robot policies from human interventions in latent space. FlowDAgger outperforms supervised fine-tuning and latent-space RL baselines and preserves pretrained skills on held-out tasks, offering a practical path for adapting robot foundation models in the real world.

Abstract

Pretrained generative robot policies based on flow matching and diffusion have achieved impressive results across a wide range of manipulation tasks. Yet real-world deployments routinely expose failure modes outside the pretraining distribution. Closing these gaps typically requires large-scale data collection or online reinforcement learning on physical hardware, which is impractical for rapid and safe adaptation. We present FlowDAgger, a sample- and compute-efficient method for adapting frozen generative robot policies from human interventions in latent space. Our key idea is action inversion: each human expert action is mapped to the noise that would have produced it under the frozen base policy, using reverse-time integration followed by local refinement. The resulting inverted noise provides supervision for a lightweight latent policy that steers the base model at deployment time, enabling rapid skill acquisition while preserving its behavioral priors. We evaluate FlowDAgger in simulation and on real-world bimanual and single-arm manipulation, adapting both action-head VLAs and world-action models from a handful of interventions. FlowDAgger outperforms supervised fine-tuning and latent-space RL baselines and preserves pretrained skills on held-out tasks, offering a practical path for adapting robot foundation models in the real world. Website: https://microsoft.github.io/FlowDAgger

Affiliations

Not identified