2025
InfoPO: On Mutual Information Maximization for Large Language Model Alignment
NAACL 2025long
We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward models and online sampling. Despite these benefits, these meth…