Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
Conventional preference learning methods often prioritize opinions held more widely when aggregating preferences from multiple evaluators. This may result in policies that are biased in favor of some types of opinions or groups and susceptible to strategic manipulation. To address this issue, we de…