← Search

Aleksandr Podkopaev

7 accepted papers

2026

Dependence-Aware Label Aggregation for LLM-as-a-Judge via Ising Models

ICML 2026poster

Large-scale AI evaluation increasingly relies on aggregating binary judgments from $K$ annotators, including LLMs used as judges. Most classical methods, e.g., Dawid-Skene or (weighted) majority voting, assume annotators are conditionally independent given the true label $Y\in\\{0,1\\}$, an assumpti…

Cited by 0SourceScholar
2023

Sequential Kernelized Independence Testing

ICML 2023poster

Independence testing is a classical statistical problem that has been extensively studied in the batch setting when one fixes the sample size before collecting data. However, practitioners often prefer procedures that adapt to the complexity of a problem at hand instead of setting sample size in adv…

Cited by 25SourcePDFScholar
2022

Tracking the risk of a deployed model and detecting harmful distribution shifts

ICLR 2022poster

When deployed in the real world, machine learning models inevitably encounter changes in the data distribution, and certain---but not all---distribution shifts could result in significant performance degradation. In practice, it may make sense to ignore benign shifts, under which the performance of…

Cited by 27SourcePDFScholar
2021

Distribution-free uncertainty quantification for classification under label shift

UAI 2021poster

Trustworthy deployment of ML models requires a proper measure of uncertainty, especially in safety-critical applications. We focus on uncertainty quantification (UQ) for classification problems via two avenues — prediction sets using conformal prediction and calibration of probabilistic predictors b…

Cited by 109SourcePDFScholar
2020

Distribution-free binary classification: prediction sets, confidence intervals and calibration

NeurIPS 2020spotlight

We study three notions of uncertainty quantification---calibration, confidence intervals and prediction sets---for binary classification in the distribution-free setting, that is without making any distributional assumptions on the data. With a focus towards calibration, we establish a 'tripod' of t…

Cited by 103SourcePDFScholar