RESEARCH PAPERYear 2024

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal; Nino Vieillard; Yongchao Zhou; Piotr Stanczyk; Sabela Ramos Garea; Matthieu Geist; Olivier Bachem

Classification

View four quadrants
Major category
Foundational work
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Explicit in survey

Category review. GKD develops general on-policy teacher/student token-distribution training for language models. Its student produces text, and the method is foundational post-training/distillation background rather than a language backbone or world/action module. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX