RESEARCH PAPERYear 2017
Deep Reinforcement Learning from Human Preferences
Classification
View four quadrants- Major category
- Foundational work
- Quadrant
- Not applicable
- Architecture
- Not applicable
- Prediction paradigm
- Not applicable
- Subcategories
- Training optimization & distillation
- Source review status
- Explicit in survey
Category review. This 2017 preference-learning foundation couples human clip comparisons, a learned scalar reward and model-free policy optimization. It supplies historical RLHF methodology for later post-training, not future-state modeling or a language/vision/action representation module. Reading evidence
Contribution
A contribution summary has not been added yet.
Abstract
An abstract has not been added yet.
Affiliations
Not listed in the collection.