2026
Adaptive KL Control for Direct Preference Optimization in Instruction-Following LLMs
AAAI 2026technical
The scaling parameter β in Direct Preference Optimization governs a fundamental trade-off: low β produces weak gradients that fail to learn from ambiguous preferences, while high β amplifies updates and causes excessive drift from the reference policy. Prior work treats β as fixed or scheduled throu