← Search

David Sontag

51 accepted papers

2026

Reinforcement Learning with Evolving Rubrics for Deep Research

ICML 2026oral

Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via reinforcement learning with verifiable rewards, which does not extend to realistic long-form tasks. We addr…

Cited by 0SourceScholar
2026

Uncovering Bias Mechanisms in Observational Studies

ICML 2026poster

Observational studies are a key resource for causal inference but are often affected by systematic biases. Prior work has focused mainly on detecting these biases, via sensitivity analyses and comparisons with randomized controlled trials, or mitigating them through debiasing techniques. However, th…

Cited by 0SourceScholar
2024

Benchmarking Observational Studies with Experimental Data under Right-Censoring

AISTATS 2024poster

Drawing causal inferences from observational studies (OS) requires unverifiable validity assumptions; however, one can falsify those assumptions by benchmarking the OS with experimental data from a randomized controlled trial (RCT). A major limitation of existing procedures is not accounting for cen…

2024

Learning to Decode Collaboratively with Multiple Language Models

ACL 2024long

We propose a method to teach multiple large language models (LLM) to collaborate by interleaving their generations at the token level. We model the decision of which LLM generates the next token as a latent variable. By optimizing the marginal likelihood of a training set under our latent variable m…

2024

Med-Real2Sim: Non-Invasive Medical Digital Twins using Physics-Informed Self-Supervised Learning

NeurIPS 2024poster

A digital twin is a virtual replica of a real-world physical phenomena that uses mathematical modeling to characterize and simulate its defining features. By constructing digital twins for disease processes, we can perform in-silico simulations that mimic patients' health conditions and counterfactu…

2024

Prediction-powered Generalization of Causal Inferences

ICML 2024poster

Causal inferences from a randomized controlled trial (RCT) may not pertain to a *target* population where some effect modifiers have a different distribution. Prior work studies *generalizing* the results of a trial to a target population with no outcome but covariate data available. We show how the…

2023

Effective Human-AI Teams via Learned Natural Language Rules and Onboarding

NeurIPS 2023spotlight

People are relying on AI agents to assist them with various tasks. The human must know when to rely on the agent, collaborate with the agent, or ignore its suggestions. In this work, we propose to learn rules grounded in data regions and described in natural language that illustrate how the human sh…

2023

Falsification of Internal and External Validity in Observational Studies via Conditional Moment Restrictions

AISTATS 2023poster

Randomized Controlled Trials (RCT)s are relied upon to assess new treatments, but suffer from limited power to guide personalized treatment decisions. On the other hand, observational (i.e., non-experimental) studies have large and diverse populations, but are prone to various biases (e.g. residual…

Cited by 10SourcePDFScholar
2023

TabLLM: Few-shot Classification of Tabular Data with Large Language Models

AISTATS 2023poster

We study the application of large language models to zero-shot and few-shot classification of tabular data. We prompt the large language model with a serialization of the tabular data to a natural-language string, together with a short description of the classification problem. In the few-shot setti…

2023

Who Should Predict? Exact Algorithms For Learning to Defer to Humans

AISTATS 2023poster

Automated AI classifiers should be able to defer the prediction to a human decision maker to ensure more accurate predictions. In this work, we jointly train a classifier with a rejector, which decides on each data point whether the classifier or the human should predict. We show that prior approach…

2022

Clustering Interval-Censored Time-Series for Disease Phenotyping

AAAI 2022technical

Unsupervised learning is often used to uncover clusters in data. However, different kinds of noise may impede the discovery of useful patterns from real-world time-series data. In this work, we focus on mitigating the interference of interval censoring in the task of clustering for disease phenotypi…

2022

Co-training Improves Prompt-based Learning for Large Language Models

ICML 2022spotlight

We demonstrate that co-training (Blum & Mitchell, 1998) can improve the performance of prompt-based learning by using unlabeled data. While prompting has emerged as a promising paradigm for few-shot and zero-shot learning, it is often brittle and requires much larger models compared to the standard…

2022

ETAB: A Benchmark Suite for Visual Representation Learning in Echocardiography

NeurIPS 2022accept

Echocardiography is one of the most commonly used diagnostic imaging modalities in cardiology. Application of deep learning models to echocardiograms can enable automated identification of cardiac structures, estimation of cardiac function, and prediction of clinical outcomes. However, a major hindr…

Cited by 4SourcePDFScholar
2022

Evaluating Robustness to Dataset Shift via Parametric Robustness Sets

NeurIPS 2022accept

We give a method for proactively identifying small, plausible shifts in distribution which lead to large differences in model performance. These shifts are defined via parametric changes in the causal mechanisms of observed variables, where constraints on parameters yield a "robustness set" of plau…

2022

Falsification before Extrapolation in Causal Effect Estimation

NeurIPS 2022accept

Randomized Controlled Trials (RCTs) represent a gold standard when developing policy guidelines. However, RCTs are often narrow, and lack data on broader populations of interest. Causal effects in these populations are often estimated using observational datasets, which may suffer from unobserved c…

2022

Large language models are few-shot clinical information extractors

EMNLP 2022main

A long-running goal of the clinical NLP community is the extraction of important variables trapped in clinical notes. However, roadblocks have included dataset shift from the general domain and a lack of public clinical corpora and annotations. In this work, we show that large language models, such…

Cited by 448SourcePDFScholar
2022

Leveraging Time Irreversibility with Order-Contrastive Pre-training

AISTATS 2022poster

Label-scarce, high-dimensional domains such as healthcare present a challenge for modern machine learning techniques. To overcome the difficulties posed by a lack of labeled data, we explore an "order-contrastive" method for self-supervised pre-training on longitudinal data. We sample pairs of time…

Cited by 13SourcePDFScholar
2022

Sample Efficient Learning of Predictors that Complement Humans

ICML 2022spotlight

One of the goals of learning algorithms is to complement and reduce the burden on human decision makers. The expert deferral setting wherein an algorithm can either predict on its own or defer the decision to a downstream expert helps accomplish this goal. A fundamental aspect of this setting is the…

2022

Teaching Humans When to Defer to a Classifier via Exemplars

AAAI 2022technical

Expert decision makers are starting to rely on data-driven automated agents to assist them with various tasks. For this collaboration to perform properly, the human decision maker must have a mental model of when and when not to rely on the agent. In this work, we aim to ensure that human decision m…

2022

Using time-series privileged information for provably efficient learning of prediction models

AISTATS 2022poster

We study prediction of future outcomes with supervised models that use privileged information during learning. The privileged information comprises samples of time series observed between the baseline time of prediction and the future outcome; this information is only available at training time whic…

2021

Beyond Perturbation Stability: LP Recovery Guarantees for MAP Inference on Noisy Stable Instances

AISTATS 2021poster

Several works have shown that perturbation stable instances of the MAP inference problem can be solved exactly using a natural linear programming (LP) relaxation. However, most of these works give few (or no) guarantees for the LP solutions on instances that do not satisfy the relatively strict pert…

Cited by 4SourcePDFScholar
2021

CLIP: A Dataset for Extracting Action Items for Physicians from Hospital Discharge Notes

ACL 2021long

Continuity of care is crucial to ensuring positive health outcomes for patients discharged from an inpatient hospital setting, and improved information sharing can help. To share information, caregivers write discharge notes containing action items to share with patients and their future caregivers,…

2021

Deep Contextual Clinical Prediction with Reverse Distillation

AAAI 2021technical

Healthcare providers are increasingly using machine learning to predict patient outcomes to make meaningful interventions. However, despite innovations in this area, deep learning models often struggle to match performance of shallow linear models in predicting these outcomes, making it difficult to…

2021

Finding Regions of Heterogeneity in Decision-Making via Expected Conditional Covariance

NeurIPS 2021poster

Individuals often make different decisions when faced with the same context, due to personal preferences and background. For instance, judges may vary in their leniency towards certain drug-related offenses, and doctors may vary in their preference for how to start treatment for certain types of pa…

2021

Graph Cuts Always Find a Global Optimum for Potts Models (With a Catch)

ICML 2021oral

We prove that the alpha-expansion algorithm for MAP inference always returns a globally optimal assignment for Markov Random Fields with Potts pairwise potentials, with a catch: the returned assignment is only guaranteed to be optimal for an instance within a small perturbation of the original probl…

Cited by 2SourcePDFScholar
2021

PClean: Bayesian Data Cleaning at Scale with Domain-Specific Probabilistic Programming

AISTATS 2021poster

Data cleaning is naturally framed as probabilistic inference in a generative model of ground-truth data and likely errors, but the diversity of real-world error patterns and the hardness of inference make Bayesian approaches difficult to automate. We present PClean, a probabilistic programming langu…

2021

Regularizing towards Causal Invariance: Linear Models with Proxies

ICML 2021spotlight

We propose a method for learning linear models whose predictive performance is robust to causal interventions on unobserved variables, when noisy proxies of those variables are available. Our approach takes the form of a regularization term that trades off between in-distribution performance and rob…

2020

Characterization of Overlap in Observational Studies

AISTATS 2020poster

Overlap between treatment groups is required for non-parametric estimation of causal effects. If a subgroup of subjects always receives the same intervention, we cannot estimate the effect of intervention changes on that subgroup without further assumptions. When overlap does not hold globally, ch…

2020

Empirical Study of the Benefits of Overparameterization in Learning Latent Variable Models

ICML 2020poster

One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e. training a very large model) to improving the optimization landscape of a problem, with minimal effect on statistical performance (i.e. generalization). In contrast, unsupervised…

2020

Estimation of Bounds on Potential Outcomes For Decision Making

ICML 2020poster

Estimation of individual treatment effects is commonly used as the basis for contextual decision making in fields such as healthcare, education, and economics. However, it is often sufficient for the decision maker to have estimates of upper and lower bounds on the potential outcomes of decision alt…

2019

Counterfactual Off-Policy Evaluation with Gumbel-Max Structural Causal Models

ICML 2019oral

We introduce an off-policy evaluation procedure for highlighting episodes where applying a reinforcement learned (RL) policy is likely to have produced a substantially different outcome than the observed policy. In particular, we introduce a class of structural causal models (SCMs) for generating co…

2019

Overcomplete Independent Component Analysis via SDP

AISTATS 2019poster

We present a novel algorithm for overcomplete independent components analysis (ICA), where the number of latent sources k exceeds the dimension p of observed variables. Previous algorithms either suffer from high computational complexity or make strong assumptions about the form of the mixing matrix…

Cited by 28SourcePDFScholar
2019

Support and Invertibility in Domain-Invariant Representations

AISTATS 2019poster

Learning domain-invariant representations has become a popular approach to unsupervised domain adaptation and is often justified by invoking a particular suite of theoretical results. We argue that there are two significant flaws in such arguments. First, the results in question hold only for a fixe…

2018

Optimality of Approximate Inference Algorithms on Stable Instances

AISTATS 2018poster

Approximate algorithms for structured prediction problems—such as LP relaxations and the popular α-expansion algorithm (Boykov et al. 2001)—typically far exceed their theoretical performance guarantees on real-world instances. These algorithms often find solutions that are very close to optimal. The…

Cited by 0SourcePDFScholar
2018

Semi-Amortized Variational Autoencoders

ICML 2018oral

Amortized variational inference (AVI) replaces instance-specific local inference with a global inference network. While AVI has enabled efficient training of deep generative models such as variational autoencoders (VAE), recent empirical work suggests that inference networks can produce suboptimal v…

2017

Causal Effect Inference with Deep Latent-Variable Models

NeurIPS 2017poster

Learning individual-level causal effects from observational data, such as inferring the most effective medication for a specific patient, is a problem of growing importance for policy makers. The most important aspect of inferring causal effects from observational data is the handling of confounders…

Cited by 972SourcePDFScholar
2017

Estimating individual treatment effect: generalization bounds and algorithms

ICML 2017poster

There is intense interest in applying machine learning to problems of causal inference in fields such as healthcare, economics and education. In particular, individual-level causal inference has important applications such as precision medicine. We give a new theoretical analysis and family of algor…

2017

Simultaneous Learning of Trees and Representations for Extreme Classification and Density Estimation

ICML 2017poster

We consider multi-class classification where the predictor has a hierarchical structure that allows for a very large number of labels both at train and test time. The predictive power of such models can heavily depend on the structure of the tree, and although past work showed how to learn the tree…

2016

Train and Test Tightness of LP Relaxations in Structured Prediction

ICML 2016poster

Structured prediction is used in areas such as computer vision and natural language processing to predict structured outputs such as segmentations or parse trees. In these settings, prediction is performed by MAP inference or, equivalently, by solving an integer linear program. Because of the comple…

Cited by 19SourcePDFScholar
2015

A Fast Variational Approach for Learning Markov Random Field Language Models

ICML 2015poster

Language modelling is a fundamental building block of natural language processing. However, in practice the size of the vocabulary limits the distributions applicable for this task: specifically, one has to either resort to local optimization methods, such as those used in neural language models, or…

Cited by 26SourcePDFScholar