2025
Preference Distillation via Value based Reinforcement Learning
NeurIPS 2025poster
Direct Preference Optimization (DPO) is a powerful paradigm to align language models with human preferences using pairwise comparisons. However, its binary win-or-loss supervision often proves insufficient for training small models with limited capacity. Prior works attempt to distill information f…