← Search

Rediet Abebe

8 accepted papers

2025

Direct Alignment with Heterogeneous Preferences

NeurIPS 2025poster

Alignment with human preferences is commonly framed using a universal reward function, even though human preferences are inherently heterogeneous. We formalize this heterogeneity by introducing user types and examine the limits of the homogeneity assumption. We show that aligning to heterogeneous pr…

Cited by 0SourcecodeScholar
2025

Lawma: The Power of Specialization for Legal Annotation

ICLR 2025poster

Annotation and classification of legal text are central components of empirical legal research. Traditionally, these tasks are often delegated to trained research assistants. Motivated by the advances in language modeling, empirical legal scholars are increasingly turning to commercial models, hopin…

2023

When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks

EMNLP 2023long main

Though majority vote among annotators is typically used for ground truth labels in machine learning, annotator disagreement in tasks such as hate speech detection may reflect systematic differences in opinion across groups, not noise. Thus, a crucial problem in hate speech detection is determining i…

Cited by 0SourceScholar