← Search

Muhan Lin

2 accepted papers

2024

Navigating Noisy Feedback: Enhancing Reinforcement Learning with Error-Prone Language Models

EMNLP 2024finding

The correct specification of reward models is a well-known challenge in reinforcement learning.Hand-crafted reward functions often lead to inefficient or suboptimal policies and may not be aligned with user values.Reinforcement learning from human feedback is a successful technique that can mitigate…