2026
Conditional Equivalence of DPO and RLHF: Assumptions, Failure Modes, and Provable Alignment
ICML 2026spotlight
Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implementation. We prove this equivalence is _conditional_ rather than universal, depending on an implicit assumption frequent…