2025
HelpSteer2-Preference: Complementing Ratings with Preferences
ICLR 2025poster
Reward models are critical for aligning models to follow instructions, and are typically trained following one of two popular paradigms: Bradley-Terry style or Regression style. However, there is a lack of evidence that either approach is better than the other, when adequately matched for data. This…