← Search

Daniel Chechelnitsky

1 accepted papers

2025

Rejected Dialects: Biases Against African American Language in Reward Models

NAACL 2025findings

Preference alignment via reward models helps build safe, helpful, and reliable large language models (LLMs). However, subjectivity in preference judgments and the lack of representative sampling in preference data collection can introduce new biases, hindering reward models’ fairness and equity. In…