Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization
Direct Preference Optimization (DPO) has emerged as a powerful framework for aligning large language models (LLMs) with human preferences via pairwise comparisons. However, its performance is highly sensitive to the quality of training samples: when the reference policy is poorly aligned with human…