← Search

Nikita Surnachev

1 accepted papers

2025

Learn Your Reference Model for Real Good Alignment

ICLR 2025poster

Despite the fact that offline methods for Large Language Models (LLMs) alignment do not require a direct reward model, they remain susceptible to overoptimization. This issue arises when the trained model deviates excessively from the reference policy, leading to a decrease in sample quality. We pro…

Cited by 26SourcePDFScholar