← Search

Ian Maksimov

3 accepted papers

2025

Learn Your Reference Model for Real Good Alignment

ICLR 2025poster

Despite the fact that offline methods for Large Language Models (LLMs) alignment do not require a direct reward model, they remain susceptible to overoptimization. This issue arises when the trained model deviates excessively from the reference policy, leading to a decrease in sample quality. We pro…

Cited by 26SourcePDFScholar
2025

Revisiting Non-Acyclic GFlowNets in Discrete Environments

ICML 2025poster

Generative Flow Networks (GFlowNets) are a family of generative models that learn to sample objects from a given probability distribution, potentially known up to a normalizing constant. Instead of working in the object space, GFlowNets proceed by sampling trajectories in an appropriately constructe…