← Search

Pratik Patil

12 accepted papers

2026

Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning

ICML 2026poster

Self-distillation (SD), retraining a student on a mixture of ground-truth labels and a teacher’s own predictions using the same architecture and training data, often improves generalization empirically, but it is unclear when improvement is guaranteed. We study SD for ridge regression with an uncons…

Cited by 0SourceScholar
2026

Precise Asymptotics of Bagging Regularized M-estimators

ICML 2026poster

We characterize the squared prediction risk of ensemble estimators obtained through subagging (subsample bootstrap aggregating) regularized M-estimators and construct a consistent estimator for the risk. Specifically, we consider a heterogeneous collection of M?1 regularized M-estimators, each train…

Cited by 0SourceScholar
2024

"A Framework for Efficient Model Evaluation through Stratification, Sampling, and Estimation"

ECCV 2024poster

"Model performance evaluation is a critical and expensive task in machine learning and computer vision. Without clear guidelines, practitioners often estimate model accuracy using a one-time completely random selection of the data. However, by employing tailored sampling and estimation strategies, o…

2024

Asymptotically Free Sketched Ridge Ensembles: Risks, Cross-Validation, and Tuning

ICLR 2024spotlight

We employ random matrix theory to establish consistency of generalized cross validation (GCV) for estimating prediction risks of sketched ridge regression ensembles, enabling efficient and consistent tuning of regularization and sketching parameters. Our results hold for a broad class of asymptotica…

2024

Failures and Successes of Cross-Validation for Early-Stopped Gradient Descent

AISTATS 2024poster

We analyze the statistical properties of generalized cross-validation (GCV) and leave-one-out cross-validation (LOOCV) applied to early-stopped gradient descent (GD) in high-dimensional least squares regression. We prove that GCV is generically inconsistent as an estimator of the prediction risk of…

Cited by 5SourcePDFScholar
2024

Optimal Ridge Regularization for Out-of-Distribution Prediction

ICML 2024spotlight

We study the behavior of optimal ridge regularization and optimal ridge risk for out-of-distribution prediction, where the test distribution deviates arbitrarily from the train distribution. We establish general conditions that determine the sign of the optimal regularization level under covariate a…

2024

Precise Model Benchmarking with Only a Few Observations

EMNLP 2024main

How can we precisely estimate a large language model’s (LLM) accuracy on questions belonging to a specific topic within a larger question-answering dataset? The standard direct estimator, which averages the model’s accuracy on the questions in each subgroup, may exhibit high variance for subgroups (…

Cited by 1SourcePDFScholar
2023

Subsample Ridge Ensembles: Equivalences and Generalized Cross-Validation

ICML 2023oral

We study subsampling-based ridge ensembles in the proportional asymptotics regime, where the feature size grows proportionally with the sample size such that their ratio converges to a constant. By analyzing the squared prediction risk of ridge ensembles as a function of the explicit penalty $\lambd…

Cited by 13SourcePDFScholar
2022

Estimating Functionals of the Out-of-Sample Error Distribution in High-Dimensional Ridge Regression

AISTATS 2022poster

We study the problem of estimating the distribution of the out-of-sample prediction error associated with ridge regression. In contrast, the traditional object of study is the uncentered second moment of this distribution (the mean squared prediction error), which can be estimated using cross-valida…

Cited by 15SourcePDFScholar
2021

Uniform Consistency of Cross-Validation Estimators for High-Dimensional Ridge Regression

AISTATS 2021poster

We examine generalized and leave-one-out cross-validation for ridge regression in a proportional asymptotic framework where the dimension of the feature space grows proportionally with the number of observations. Given i.i.d. samples from a linear model with an arbitrary feature covariance and a sig…

Cited by 64SourcePDFScholar