← Search

Paul F Christiano

2 accepted papers

2020

Learning to summarize with human feedback

NeurIPS 2020poster

As language models become more powerful, training and evaluation are increasingly bottlenecked by the data and metrics used for a particular task. For example, summarization models are often trained to predict human reference summaries and evaluated using ROUGE, but both of these metrics are rough…

2017

Deep Reinforcement Learning from Human Preferences

NeurIPS 2017poster

For sophisticated reinforcement learning (RL) systems to interact usefully with real-world environments, we need to communicate complex goals to these systems. In this work, we explore goals defined in terms of (non-expert) human preferences between pairs of trajectory segments. Our approach separat…

Cited by 4197SourcePDFScholar