RESEARCH PAPERYear 2024
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
Classification
View four quadrants- Major category
- Foundational work
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Training optimization & distillation
- Source review status
- Explicit in survey
Category review. GKD develops general on-policy teacher/student token-distribution training for language models. Its student produces text, and the method is foundational post-training/distillation background rather than a language backbone or world/action module. Reading evidence
Contribution
A contribution summary has not been added yet.
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.