2026
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
AAAI 2026technical
Uniform-reward reinforcement learning from human feedback (RLHF), which trains a single reward model to represent the preferences of all annotators, fails to capture the diversity of opinions across sub-populations, inadvertently favoring dominant groups. The state-of-the-art, MaxMin-RLHF, addresses