← Search

Tyna Eloundou

2 accepted papers

2025

First-Person Fairness in Chatbots

ICLR 2025spotlight

Evaluating chatbot fairness is crucial given their rapid proliferation, yet typical chatbot tasks (e.g., resume writing, entertainment) diverge from the institutional decision-making tasks (e.g., resume screening) which have traditionally been central to discussion of algorithmic fairness. The open-…

Cited by 8SourcePDFScholar
2025

SEAL: Systematic Error Analysis for Value ALignment

AAAI 2025technical

Reinforcement Learning from Human Feedback (RLHF) aligns language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base models. Despite its importance, the internal mechanisms of RLHF remain poorly understood. This paper introd…