2025
Alleviating Shifted Distribution in Human Preference Alignment through Meta-Learning
AAAI 2025technical
The capability of the reward model (RM) is crucial for the success of Reinforcement Learning from Human Feedback (RLHF) in aligning with human preferences. However, as training progresses, the output space distribution of the policy model shifts. The RM, initially trained on responses sampled from t…