2026
The Secret Engine Behind RLHF: It's Contarstive Learning All Along
ICML 2026poster
Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that…