2025
Not All Voices Are Rewarded Equally: Probing and Repairing Reward Models across Human Diversity
EMNLP 2025
The advancement of Large Language Models (LLMs) has made ensuring their trustworthiness increasingly critical, especially in terms of fairness across diverse human groups. While modern LLMs are aligned with user preferences through Reinforcement Learning from Human Feedback (RLHF), the reward models