NAACL 2025long0 citations

InfoPO: On Mutual Information Maximization for Large Language Model Alignment

Teng Xiao, Zhen Ge, Sujay Sanghavi, Tian Wang, Julian Katz-Samuels, Marc Versage, Qingjun Cui, Trishul Chilimbi

Abstract

We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward models and online sampling. Despite these benefits, these methods rely on explicit assumptions about the Bradley-Terry (BT) model, which makes them prone to overfitting and results in suboptimal performance, particularly on reasoning-heavy tasks. To address these challenges, we propose a principled preference fine-tuning algorithm called InfoPO, which effectively and efficiently aligns large language models using preference data. InfoPO eliminates the reliance on the BT model and prevents the likelihood of the chosen response from decreasing. Extensive experiments confirm that InfoPO consistently outperforms established baselines on widely used open benchmarks, particularly in reasoning tasks.

BibTeX
@inproceedings{xiao-etal-2025-infopo,
    title = "{I}nfo{PO}: On Mutual Information Maximization for Large Language Model Alignment",
    author = "Xiao, Teng  and
      Ge, Zhen  and
      Sanghavi, Sujay  and
      Wang, Tian  and
      Katz-Samuels, Julian  and
      Versage, Marc  and
      Cui, Qingjun  and
      Chilimbi, Trishul",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.585/",
    pages = "11699--11711",
    ISBN = "979-8-89176-189-6"
}
InfoPO: On Mutual Information Maximization for Large Language Model Alignment · NAACL 2025