ICLR 2025poster26 citations

BOND: Aligning LLMs with Best-of-N Distillation

Pier Giuseppe Sessa, Robert Dadashi-Tazehozi, Leonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Rame, Bobak Shahriari, Sarah Perrin

Abstract

Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-time strategy is Best-of-N sampling that selects the best generation among N candidates. In this paper, we propose Best-of-N Distillation (BOND), a novel RLHF algorithm that seeks to emulate Best-of-N but without its significant computational overhead at inference time. Specifically, BOND is a distribution matching algorithm that forces the distribution of generations from the policy to get closer to the Best-of-N distribution. We use the Jeffreys divergence (a linear combination of forward and backward KL) to balance between mode-covering and mode-seeking behavior, and derive an iterative formulation that utilizes a moving anchor for efficiency. We demonstrate the effectiveness of our approach and several design choices through experiments on abstractive summarization and Gemma models.

LLMAlignmentRLHFBest-of-N
BibTeX
@inproceedings{
sessa2025bond,
title={{BOND}: Aligning {LLM}s with Best-of-N Distillation},
author={Pier Giuseppe Sessa and Robert Dadashi-Tazehozi and Leonard Hussenot and Johan Ferret and Nino Vieillard and Alexandre Rame and Bobak Shahriari and Sarah Perrin and Abram L. Friesen and Geoffrey Cideron and Sertan Girgin and Piotr Stanczyk and Andrea Michi and Danila Sinopalnikov and Sabela Ramos Garea and Am{\'e}lie H{\'e}liou and Aliaksei Severyn and Matthew Hoffman and Nikola Momchev and Olivier Bachem},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=0tAXMiSufG}
}
BOND: Aligning LLMs with Best-of-N Distillation · ICLR 2025