2026
The Viscosity of Logic: Phase Transitions and Hysteresis in DPO Alignment
ICML 2026poster
Direct Preference Optimization (DPO) is often tuned as if increasing alignment pressure (controlled by $\beta$) yields progressively “better” behavior. We instead treat $\beta$ as a control parameter and densely sweep it for three 7B open-weight families under a fixed DPO recipe. In Mistral, capabil…