← Search

Amélie Héliou

4 accepted papers

2025

BOND: Aligning LLMs with Best-of-N Distillation

ICLR 2025poster

Reinforcement learning from human feedback (RLHF) is a key driver of quality and safety in state-of-the-art large language models. Yet, a surprisingly simple and strong inference-time strategy is Best-of-N sampling that selects the best generation among N candidates. In this paper, we propose Best-o…

Cited by 26SourcePDFScholar
2021

Zeroth-Order Non-Convex Learning via Hierarchical Dual Averaging

ICML 2021spotlight

We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization {–} i.e., learning processes where, at each stage, the optimizer is facing an unknown non-convex loss function and only receives the incurred loss as feedback. The proposed class of policies relies on…

Cited by 16SourcePDFScholar
2020

Gradient-free Online Learning in Continuous Games with Delayed Rewards

ICML 2020poster

Motivated by applications to online advertising and recommender systems, we consider a game-theoretic model with delayed rewards and asynchronous, payoff-based feedback. In contrast to previous work on delayed multi-armed bandits, we focus on games with continuous action spaces, and we examine the l…

Cited by 50SourcePDFScholar
2020

Online Non-Convex Optimization with Imperfect Feedback

NeurIPS 2020poster

We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes – or otherwise constructs – an inexact model for the loss function encountered at each stage, and we propose a mixed-strategy learning policy based on dual averaging. In this…

Cited by 26SourcePDFScholar