RESEARCH PAPERYear 2020
Learning to summarize with human feedback
Classification
View four quadrants- Major category
- Foundational work
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Training optimization & distillation
- Source review status
- Explicit in survey
Category review. This 2020 work establishes human-preference reward learning and PPO-based language-model post-training. Its output is a summarization policy and learned reward, making it historical post-training background rather than a WAM model component. Reading evidence
Contribution
A contribution summary has not been added yet.
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.