← Search

Matteo Cargnelutti

1 accepted papers

2025

SEAL: Systematic Error Analysis for Value ALignment

AAAI 2025technical

Reinforcement Learning from Human Feedback (RLHF) aligns language models (LMs) with human values by training reward models (RMs) on binary preferences and using these RMs to fine-tune the base models. Despite its importance, the internal mechanisms of RLHF remain poorly understood. This paper introd…