Layer-wise Gradient Disentanglement: Decoupling Semantics and Preferences in Direct Preference Optimization
Direct Preference Optimization (DPO) has become the dominant approach for aligning large language models with human preferences. However, standard DPO treats all preference pairs uniformly, overlooking the heterogeneous nature of the learning problem: some samples demand sophisticated semantic under…