2024
Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms
NeurIPS 2024poster
Reinforcement Learning from Human Feedback (RLHF)has been crucial to the recent success of Large Language Models (LLMs), however it is often a complex and brittle process. In the classical RLHF framework, a reward model is first trained to represent human preferences, which is in turn used by an onl…