← Search

Gholamali Aminian

10 accepted papers

2026

Best-of-N through the Smoothing Lens: KL Divergence and Regret Analysis

ICLR 2026poster

A simple yet effective method for inference-time alignment of generative models is Best-of-$N$ (BoN), where $N$ outcomes are sampled from a reference policy, evaluated using a proxy-reward model, and the highest-scoring one is selected. While prior work argues that BoN is almost optimal in reward…

Cited by 0SourceScholar
2025

Generalization and Robustness of the Tilted Empirical Risk

ICML 2025poster

The generalization error (risk) of a supervised statistical learning algorithm quantifies its prediction ability on previously unseen data. Inspired by exponential tilting, Li et al. (2021) proposed the {\it tilted empirical risk} (TER) as a non-linear risk metric for machine learning applications…

Cited by 0SourcePDFScholar
2025

KL-Regularized RLHF with Multiple Reference Models: Exact Solutions and Sample Complexity

NeurIPS 2025poster

Recent methods for aligning large language models (LLMs) with human feedback predominantly rely on a single reference model, which limits diversity, model overfitting, and underutilizes the wide range of available pre-trained models. Incorporating multiple reference models has the potential to addre…

Cited by 0SourceScholar
2025

Log-Sum-Exponential Estimator for Off-Policy Evaluation and Learning

ICML 2025spotlight

Off-policy learning and evaluation leverage logged bandit feedback datasets, which contain context, action, propensity score, and feedback for each data point. These scenarios face significant challenges due to high variance and poor performance with low-quality propensity scores and heavy-tailed re…

2025

Pessimistic Data Integration for Policy Evaluation

NeurIPS 2025poster

This paper studies how to integrate historical control data with experimental data to enhance A/B testing, while addressing the distributional shift between historical and experimental datasets. We propose a pessimistic data integration method that combines two causal effect estimators constructed b…

Cited by 0SourceScholar
2024

Generalization Error of Graph Neural Networks in the Mean-field Regime

ICML 2024poster

This work provides a theoretical framework for assessing the generalization error of graph neural networks in the over-parameterized regime, where the number of parameters surpasses the quantity of data points. We explore two widely utilized types of graph neural networks: graph convolutional neural…

2023

How Does Pseudo-Labeling Affect the Generalization Error of the Semi-Supervised Gibbs Algorithm?

AISTATS 2023poster

We provide an exact characterization of the expected generalization error (gen-error) for semi-supervised learning (SSL) with pseudo-labeling via the Gibbs algorithm. The gen-error is expressed in terms of the symmetrized KL information between the output hypothesis, the pseudo-labeled dataset, and…

Cited by 6SourcePDFScholar
2022

An Information-theoretical Approach to Semi-supervised Learning under Covariate-shift

AISTATS 2022poster

A common assumption in semi-supervised learning is that the labeled, unlabeled, and test data are drawn from the same distribution. However, this assumption is not satisfied in many applications. In many scenarios, the data is collected sequentially (e.g., healthcare) and the distribution of the dat…

Cited by 32SourcePDFScholar
2022

Characterizing and Understanding the Generalization Error of Transfer Learning with Gibbs Algorithm

AISTATS 2022poster

We provide an information-theoretic analysis of the generalization ability of Gibbs-based transfer learning algorithms by focusing on two popular empirical risk minimization (ERM) approaches for transfer learning, $\alpha$-weighted-ERM and two-stage-ERM. Our key result is an exact characterization o…

Cited by 17SourcePDFScholar
2021

An Exact Characterization of the Generalization Error for the Gibbs Algorithm

NeurIPS 2021poster

Various approaches have been developed to upper bound the generalization error of a supervised learning algorithm. However, existing bounds are often loose and lack of guarantees. As a result, they may fail to characterize the exact generalization ability of a learning algorithm. Our main contributi…

Cited by 54SourcePDFScholar