2025
Uncertainty-Aware Iterative Preference Optimization for Enhanced LLM Reasoning
ACL 2025long
Direct Preference Optimization (DPO) has recently emerged as an efficient and effective method for aligning large language models with human preferences. However, constructing high-quality preference datasets remains challenging, often necessitating expensive manual or powerful LM annotations. Addit…