2025
SEAL: Systematic Error Analysis for Value ALignment
AAAI 2025technical
Reinforcement Learning from Human Feedback (RLHF) aligns language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base models. Despite its importance, the internal mechanisms of RLHF remain poorly understood. This paper introd…