← Search

Zachary Lipton

25 accepted papers

2026

Expert Routing with Synthetic Data for Domain Incremental Learning

ICML 2026poster

In many real-world settings, regulations and economic incentives permit the sharing of models but not data across institutional boundaries. In such scenarios, practitioners might hope to adapt models to new domains, without losing performance on previous domains (so-called catastrophic forgetting). …

Cited by 0SourceScholar
2026

No Free Lunch: Non-Asymptotic Analysis of Prediction-Powered Inference

ICML 2026poster

Prediction-Powered Inference (PPI) is a popular strategy for combining gold-standard and possibly noisy pseudo-labels to perform statistical estimation. Prior work has shown an asymptotic \enquote{free lunch} for PPI++, an adaptive form of PPI, showing that the \textit{asymptotic} variance of PPI++ …

Cited by 0SourceScholar
2024

Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends

ACL 2024long

Recent advancements in large language models (LLMs) have significantly advanced the capabilities of summarization systems.However, they continue to face a persistent challenge: hallucination. While prior work has extensively examined LLMs in news domains, evaluation of dialogue summarization has pri…

2024

Auditing Fairness under Unobserved Confounding

AISTATS 2024poster

A fundamental problem in decision-making systems is the presence of inequity along demographic lines. However, inequity can be difficult to quantify, particularly if our notion of equity relies on hard-to-measure notions like risk (e.g., equal access to treatment for those who would die without it).…

2024

Timing as an Action: Learning When to Observe and Act

AISTATS 2024poster

In standard reinforcement learning setups, the agent receives observations and performs actions at evenly spaced intervals. However, in many real-world settings, observations are expensive, forcing agents to commit to courses of action for designated periods of time. Consider that doctors, after eac…

Cited by 3SourcePDFScholar
2023

Downstream Datasets Make Surprisingly Good Pretraining Corpora

ACL 2023long

For most natural language processing tasks, the dominant practice is to finetune large pretrained transformer models (e.g., BERT) using smaller downstream datasets. Despite the success of this approach, it remains unclear to what extent these gainsare attributable to the massive background corpora e…

2023

Risk-limiting financial audits via weighted sampling without replacement

UAI 2023poster

We introduce the notion of risk-limiting financial audits (RLFA): procedures that manually evaluate a subset of $N$ financial transactions to check the validity of a claimed assertion $\mathcal{A}$ about the transactions. More specifically, RLFA satisfy two properties: (i) if $\mathcal{A}$ is false…

Cited by 5SourcePDFScholar
2022

Off-Policy Risk Assessment for Markov Decision Processes

AISTATS 2022poster

Addressing such diverse ends as mitigating safety risks, aligning agent behavior with human preferences, and improving the efficiency of learning, an emerging line of reinforcement learning research addresses the entire distribution of returns and various risk functionals that depend upon it. In the…

Cited by 8SourcePDFScholar
2022

Supervised Learning with General Risk Functionals

ICML 2022spotlight

Standard uniform convergence results bound the generalization gap of the expected loss over a hypothesis class. The emergence of risk-sensitive learning requires generalization guarantees for functionals of the loss distribution beyond the expectation. While prior works specialize in uniform converg…

Cited by 11SourcePDFScholar
2021

Correcting Exposure Bias for Link Recommendation

ICML 2021spotlight

Link prediction methods are frequently applied in recommender systems, e.g., to suggest citations for academic papers or friends in social networks. However, exposure bias can arise when users are systematically underexposed to certain relevant items. For example, in citation networks, authors might…

2021

On Proximal Policy Optimization’s Heavy-tailed Gradients

ICML 2021spotlight

Modern policy gradient algorithms such as Proximal Policy Optimization (PPO) rely on an arsenal of heuristics, including loss clipping and gradient clipping, to ensure successful learning. These heuristics are reminiscent of techniques from robust statistics, commonly used for estimation in outlier-…

Cited by 15SourcePDFScholar
2021

RATT: Leveraging Unlabeled Data to Guarantee Generalization

ICML 2021oral

To assess generalization, machine learning scientists typically either (i) bound the generalization gap and then (after training) plug in the empirical risk to obtain a bound on the true risk; or (ii) validate empirically on holdout data. However, (i) typically yields vacuous guarantees for overpara…

2020

A Unified View of Label Shift Estimation

NeurIPS 2020poster

Under label shift, the label distribution $p(y)$ might change but the class-conditional distributions $p(x|y)$ do not. There are two dominant approaches for estimating the label marginal. BBSE, a moment-matching approach based on confusion matrices, is provably consistent and provides interpretable…

2020

Learning The Difference That Makes A Difference With Counterfactually-Augmented Data

ICLR 2020spotlight

Despite alarm over the reliance of machine learning systems on so-called spurious patterns, the term lacks coherent meaning in standard statistical frameworks. However, the language of causality offers clarity: spurious associations are due to confounding (e.g., a common cause), but not direct or in…

Cited by 661SourceScholar
2020

Uncertainty-Aware Lookahead Factor Models for Quantitative Investing

ICML 2020poster

On a periodic basis, publicly traded companies report fundamentals, financial data including revenue, earnings, debt, among others. Quantitative finance research has identified several factors, functions of the reported data that historically correlate with stock market performance. In this paper, w…

Cited by 18SourcePDFScholar
2019

Domain Adaptation with Asymmetrically-Relaxed Distribution Alignment

ICML 2019oral

Domain adaptation addresses the common situation in which the target distribution generating our test data differs from the source distribution generating our training data. While absent assumptions, domain adaptation is impossible, strict conditions, e.g. covariate or label shift, enable principled…

Cited by 172SourcePDFScholar
2019

Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift

NeurIPS 2019poster

We might hope that when faced with unexpected inputs, well-designed software systems would fire off warnings. Machine learning (ML) systems, however, which depend strongly on properties of their inputs (e.g. the i.i.d. assumption), tend to fail silently. This paper explores the problem of building M…

2019

Game Design for Eliciting Distinguishable Behavior

NeurIPS 2019poster

The ability to inferring latent psychological traits from human behavior is key to developing personalized human-interacting machine learning systems. Approaches to infer such traits range from surveys to manually-constructed experiments and games. However, these traditional games are limited becaus…

Cited by 2SourcePDFScholar
2019

Learning Robust Global Representations by Penalizing Local Predictive Power

NeurIPS 2019poster

Despite their renowned in-domain predictive power, convolutional neural networks are known to rely more on high-frequency patterns that humans deem superficial than on low-frequency patterns that agree better with intuitions about what constitutes category membership. This paper proposes a method fo…

2018

Born Again Neural Networks

ICML 2018oral

Knowledge Distillation (KD) consists of transferring “knowledge” from one machine learning model (the teacher) to another (the student). Commonly, the teacher is a high-capacity model with formidable performance, while the student is more compact. By transferring knowledge, one hopes to benefit from…

Cited by 1313SourcePDFScholar
2018

Detecting and Correcting for Label Shift with Black Box Predictors

ICML 2018oral

Faced with distribution shift between training and test set, we wish to detect and quantify the shift, and to correct our classifiers without test set labels. Motivated by medical diagnosis, where diseases (targets), cause symptoms (observations), we focus on label shift, where the label marginal p(…

2018

Does mitigating ML's impact disparity require treatment disparity?

NeurIPS 2018poster

Following precedent in employment discrimination law, two notions of disparity are widely-discussed in papers on fairness and ML. Algorithms exhibit treatment disparity if they formally treat members of protected subgroups differently; algorithms exhibit impact disparity when outcomes differ across…