Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
Direct alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective for matching relative preferences, these methods have been widely observed to depress the absolute likelihoods of example responses. Consequen…