2024
Enhancing Alignment using Curriculum Learning & Ranked Preferences
EMNLP 2024finding
Direct Preference Optimization (DPO) is an effective technique that leverages pairwise preference data (one chosen and rejected response per prompt) to align LLMs to human preferences. In practice, multiple responses could exist for a given prompt with varying quality relative to each other. We prop…