← Search

Amir R. Asadi

3 accepted papers

2026

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

ICLR 2026poster

A simple yet effective method for inference-time alignment of generative models is Best-of-$N$ (BoN), where $N$ outcomes are sampled from a reference policy, evaluated using a proxy-reward model, and the highest-scoring one is selected. While prior work argues that BoN is almost optimal in reward…

Cited by 0SourceScholar
2025

Generalization and Robustness of the Tilted Empirical Risk

ICML 2025poster

The generalization error (risk) of a supervised statistical learning algorithm quantifies its prediction ability on previously unseen data. Inspired by exponential tilting, Li et al. (2021) proposed the {\it tilted empirical risk} (TER) as a non-linear risk metric for machine learning applications…

Cited by 0SourcePDFScholar
2025

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

NeurIPS 2025poster

Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underutilizes the wide range of available pre-trained models. Incorporating multiple reference models has the potential to addre…

Cited by 0SourceScholar