← Search

Alessandro Rinaldo

24 accepted papers

2026

Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning

ICML 2026poster

Self-distillation (SD), retraining a student on a mixture of ground-truth labels and a teacher’s own predictions using the same architecture and training data, often improves generalization empirically, but it is unclear when improvement is guaranteed. We study SD for ridge regression with an uncons…

Cited by 0SourceScholar
2025

On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts

NeurIPS 2025poster

The softmax-contaminated mixture of experts (MoE) model is deployed when a large-scale pre-trained model, which plays the role of a fixed expert, is fine-tuned for learning downstream tasks by including a new contamination part, or prompt, functioning as a new, trainable expert. Despite its populari…

Cited by 0SourceScholar
2024

On the estimation of persistence intensity functions and linear representations of persistence diagrams

AISTATS 2024poster

Persistence diagrams are one of the most popular types of data summaries used in Topological Data Analysis. The prevailing statistical approach to analyzing persistence diagrams is concerned with filtering out topological noise. In this paper, we adopt a different viewpoint and aim at estimating the…

Cited by 1SourcePDFScholar
2024

Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts

NeurIPS 2024poster

The softmax gating function is arguably the most popular choice in mixture of experts modeling. Despite its widespread use in practice, the softmax gating may lead to unnecessary competition among experts, potentially causing the undesirable phenomenon of representation collapse due to its inherent…

Cited by 7SourcePDFScholar
2023

Divide and Conquer Dynamic Programming: An Almost Linear Time Change Point Detection Methodology in High Dimensions

ICML 2023poster

We develop a novel, general and computationally efficient framework, called Divide and Conquer Dynamic Programming (DCDP), for localizing change points in time series data with high-dimensional features. DCDP deploys a class of greedy algorithms that are applicable to a broad variety of high-dimensi…

2022

$\ell_∞$-Bounds of the MLE in the BTL Model under General Comparison Graphs

UAI 2022poster

The Bradley-Terry-Luce (BTL) model is a popular statistical approach for estimating the global ranking of a collection of items using pairwise comparisons. To ensure accurate ranking, it is essential to obtain precise estimates of the model parameters in the $\ell_{\infty}$-loss. The difficulty of t…

2022

Denoising and change point localisation in piecewise-constant high-dimensional regression coefficients

AISTATS 2022poster

We study the theoretical properties of the fused lasso procedure originally proposed by Tibshirani et al. (2005) in the context of a linear regression model in which the regression coefficient are totally ordered and assumed to be sparse and piecewise constant. Despite its popularity, to the best of…

Cited by 11SourcePDFScholar
2022

Detecting Abrupt Changes in Sequential Pairwise Comparison Data

NeurIPS 2022accept

The Bradley-Terry-Luce (BTL) model is a classic and very popular statistical approach for eliciting a global ranking among a collection of items using pairwise comparison data. In applications in which the comparison outcomes are observed as a time series, it is often the case that data are non-stat…

2022

Estimating Functionals of the Out-of-Sample Error Distribution in High-Dimensional Ridge Regression

AISTATS 2022poster

We study the problem of estimating the distribution of the out-of-sample prediction error associated with ridge regression. In contrast, the traditional object of study is the uncentered second moment of this distribution (the mean squared prediction error), which can be estimated using cross-valida…

Cited by 15SourcePDFScholar
2022

Generalized Results for the Existence and Consistency of the MLE in the Bradley-Terry-Luce Model

ICML 2022oral

Ranking problems based on pairwise comparisons, such as those arising in online gaming, often involve a large pool of items to order. In these situations, the gap in performance between any two items can be significant, and the smallest and largest winning probabilities can be very close to zero or…

2021

Localizing Changes in High-Dimensional Regression Models

AISTATS 2021poster

This paper addresses the problem of localizing change points in high-dimensional linear regression models with piecewise constant regression coefficients. We develop a dynamic programming approach to estimate the locations of the change points whose performance improves upon the current state-of-the…

Cited by 49SourcePDFScholar
2021

Uniform Consistency of Cross-Validation Estimators for High-Dimensional Ridge Regression

AISTATS 2021poster

We examine generalized and leave-one-out cross-validation for ridge regression in a proportional asymptotic framework where the dimension of the feature space grows proportionally with the number of observations. Given i.i.d. samples from a linear model with an arbitrary feature covariance and a sig…

Cited by 64SourcePDFScholar
2020

Nonparametric Estimation in the Dynamic Bradley-Terry Model

AISTATS 2020poster

We propose a time-varying generalization of the Bradley-Terry model that allows for nonparametric modeling of dynamic global rankings of distinct teams. We develop a novel estimator that relies on kernel smoothing to pre-process the pairwise comparisons over time and is applicable in sparse settings…

2019

Are sample means in multi-armed bandits positively or negatively biased?

NeurIPS 2019spotlight

It is well known that in stochastic multi-armed bandits (MAB), the sample mean of an arm is typically not an unbiased estimator of its true mean. In this paper, we decouple three different sources of this selection bias: adaptive \emph{sampling} of arms, adaptive \emph{stopping} of the experiment, a…

Cited by 54SourcePDFScholar
2019

Statistical Analysis of Nearest Neighbor Methods for Anomaly Detection

NeurIPS 2019poster

Nearest-neighbor (NN) procedures are well studied and widely used in both supervised and unsupervised learning problems. In this paper we are concerned with investigating the performance of NN-based methods for anomaly detection. We first show through extensive simulations that NN methods compare fa…

2019

Uniform Convergence Rate of the Kernel Density Estimator Adaptive to Intrinsic Volume Dimension

ICML 2019oral

We derive concentration inequalities for the supremum norm of the difference between a kernel density estimator (KDE) and its point-wise expectation that hold uniformly over the selection of the bandwidth and under weaker conditions on the kernel and the data generating distribution than previously…

Cited by 40SourcePDFScholar
2017

A Sharp Error Analysis for the Fused Lasso, with Application to Approximate Changepoint Screening

NeurIPS 2017poster

In the 1-dimensional multiple changepoint detection problem, we derive a new fast error rate for the fused lasso estimator, under the assumption that the mean vector has a sparse number of changepoints. This rate is seen to be suboptimal (compared to the minimax rate) by only a factor of $\log\log{n…

Cited by 77SourcePDFScholar
2016

Statistical Inference for Cluster Trees

NeurIPS 2016poster

A cluster tree provides an intuitive summary of a density function that reveals essential structure about the high-density clusters. The true cluster tree is estimated from a finite sample from an unknown true density. This paper addresses the basic question of quantifying our uncertainty by assess…

Cited by 32SourcePDFScholar
2015

Subsampling Methods for Persistent Homology

ICML 2015poster

Persistent homology is a multiscale method for analyzing the shape of sets and functions from point cloud data arising from an unknown distribution supported on those sets. When the size of the sample is large, direct computation of the persistent homology is prohibitive due to the combinatorial nat…

Cited by 151SourcePDFScholar