2026
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
ICLR 2026poster
Aligning large language models (LLMs) is a central objective of post-training, often achieved through reward modeling and reinforcement learning methods. Among these, direct preference optimization (DPO) has emerged as a widely adopted technique that fine-tunes LLMs on preferred completions over les…