2026
When Distance Distracts: Representation Distance Bias in BT-Loss for Reward Models
ICML 2026poster
Reward models are central to Large Language Model (LLM) alignment within the framework of RLHF. The standard objective used in reward modeling is the Bradley-Terry (BT) loss, which learns from pairwise data consisting of a pair of chosen and rejected responses. In this work, we analyze the per-sampl…