← Search

Annie Ulichney

1 accepted papers

2026

Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner

ICML 2026poster

While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by Golz et al. (2025) demonstrated that the *distortion* — defined…

Cited by 0SourceScholar