RESEARCH PAPERYear 2017

Deep Reinforcement Learning from Human Preferences

Paul F. Christiano; Jan Leike; Tom Brown; Miljan Martic; Shane Legg; Dario Amodei

Classification

View four quadrants
Major category
Foundational work
Architecture
Not applicable
Prediction paradigm
Not applicable
Source review status
Explicit in survey

Category review. This 2017 preference-learning foundation couples human clip comparisons, a learned scalar reward and model-free policy optimization. It supplies historical RLHF methodology for later post-training, not future-state modeling or a language/vision/action representation module. Reading evidence

AT A GLANCE

Contribution

A contribution summary has not been added yet.

Abstract

An abstract has not been added yet.

Affiliations

Not listed in the collection.

BibTeX