← Search

Stefan Feuerriegel

49 accepted papers

2026

*Rank-Learner*: Orthogonal Ranking of Treatment Effects

ICML 2026poster

Many decision-making problems require ranking individuals by their treatment effects rather than estimating the exact effect magnitudes. Examples include prioritizing patients for preventive care interventions, or ranking customers by the expected incremental impact of an advertisement. Surprisingly…

Cited by 0SourceScholar
2026

An Orthogonal Learner for Individualized Outcomes in Markov Decision Processes

ICLR 2026poster

Predicting individualized potential outcomes in sequential decision-making is central for optimizing therapeutic decisions in personalized medicine (e.g., which dosing sequence to give to a cancer patient). However, predicting potential out- comes over long horizons is notoriously difficult. Existi…

Cited by 0SourcecodeScholar
2026

Efficient and Sharp Off-Policy Learning under Unobserved Confounding

ICLR 2026poster

We develop a novel method for personalized off-policy learning in scenarios with unobserved confounding. Thereby, we address a key limitation of standard policy learning: standard policy learning assumes unconfoundedness, meaning that no unobserved factors influence both treatment assignment and out…

Cited by 0SourcecodeScholar
2026

Foundation Models for Causal Inference via Prior-Data Fitted Networks

ICLR 2026poster

Prior-data fitted networks (PFNs) have recently been proposed as a promising way to train tabular foundation models. PFNs are transformers that are pre-trained on synthetic data generated from a prespecified prior distribution and that enable Bayesian inference through in-context learning. In this p…

Cited by 30SourcecodeScholar
2026

Frequentist Consistency of Prior-Data Fitted Networks for Causal Estimation

ICML 2026poster

Foundation models based on prior-data fitted networks (PFNs) have shown strong empirical performance in causal inference by framing it as an in-context learning problem. However, it is unclear whether PFN-based causal estimators provide uncertainty quantification that is consistent with classical fr…

Cited by 0SourceScholar
2026

GDR-learners: Orthogonal Learning of Generative Models for Potential Outcomes

ICLR 2026poster

Various deep generative models have been proposed to estimate potential outcomes distributions from observational data. However, none of them have the favorable theoretical property of general Neyman-orthogonality and, associated with it, quasi-oracle efficiency and double robustness. In this paper,…

Cited by 0SourcecodeScholar
2026

IGC-Net for conditional average potential outcome estimation over time

ICLR 2026poster

Estimating potential outcomes for treatments over time based on observational data is important for personalized decision-making in medicine. However, many existing methods for this task fail to properly adjust for time-varying confounding and thus yield biased estimates. There are only a few neural…

Cited by 10SourcecodeScholar
2026

Nonparametric LLM Evaluation from Preference Data

ICML 2026poster

Evaluating the performance of large language models (LLMs) from human preference data is crucial for obtaining LLM leaderboards. However, many existing approaches either rely on restrictive parametric assumptions or lack valid uncertainty quantification when flexible machine learning methods are use…

Cited by 0SourceScholar
2026

Overlap-Adaptive Regularization for Conditional Average Treatment Effect Estimation

ICLR 2026poster

The conditional average treatment effect (CATE) is widely used in personalized medicine to inform therapeutic decisions. However, state-of-the-art methods for CATE estimation (so-called meta-learners) often perform poorly in the presence of low overlap. In this work, we introduce a new approach to t…

Cited by 0SourceScholar
2026

Overlap-weighted orthogonal meta-learner for treatment effect estimation over time

ICLR 2026poster

Estimating heterogeneous treatment effects (HTEs) in time-varying settings is particularly challenging, as the probability of observing certain treatment sequences decreases exponentially with longer prediction horizons. Thus, the observed data contain little support for many plausible treatment seq…

Cited by 0SourcecodeScholar
2026

ProbeLLM: Automating Principled Diagnosis of LLM Failures

ICML 2026poster

Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall behind. While automated probing has been enabled by dynamic test generation, existing approaches often discover isolated failure cases, lack principled con…

Cited by 0SourceScholar
2026

SurvDiff: A Diffusion Model for Generating Synthetic Data in Survival Analysis

ICML 2026spotlight

Survival analysis is a cornerstone of clinical research by modeling time-to-event outcomes such as metastasis, disease relapse, or patient death. Unlike standard tabular data, survival data often come with incomplete event information due to dropout, or loss to follow-up. This poses unique challenge…

Cited by 0SourceScholar
2025

Conformal Prediction for Causal Effects of Continuous Treatments

NeurIPS 2025poster

Uncertainty quantification of causal effects is crucial for safety-critical applications such as personalized medicine. A powerful approach for this is conformal prediction, which has several practical benefits due to model-agnostic finite-sample guarantees. Yet, existing methods for conformal predi…

Cited by 0SourcecodeScholar
2025

Constructing Confidence Intervals for Average Treatment Effects from Multiple Datasets

ICLR 2025poster

Constructing confidence intervals (CIs) for the average treatment effect (ATE) from patient records is crucial to assess the effectiveness and safety of drugs. However, patient records typically come from different hospitals, thus raising the question of how multiple observational/experimental datas…

2025

Differentially private learners for heterogeneous treatment effects

ICLR 2025poster

Patient data is widely used to estimate heterogeneous treatment effects and understand the effectiveness and safety of drugs. Yet, patient data includes highly sensitive information that must be kept private. In this work, we aim to estimate the conditional average treatment effect (CATE) from obser…

Cited by 0SourcePDFScholar
2025

Improving the Generation and Evaluation of Synthetic Data for Downstream Medical Causal Inference

NeurIPS 2025poster

Causal inference is essential for developing and evaluating medical interventions, yet real-world medical datasets are often difficult to access due to regulatory barriers. This makes synthetic data a potentially valuable asset that enables these medical analyses, along with the development of new i…

Cited by 0SourceScholar
2025

LLM-Driven Treatment Effect Estimation Under Inference Time Text Confounding

NeurIPS 2025poster

Estimating treatment effects is crucial for personalized decision-making in medicine, but this task faces unique challenges in clinical practice. At training time, models for estimating treatment effects are typically trained on well-structured medical datasets that contain detailed patient informat…

Cited by 0SourceScholar
2025

Learning Representations of Instruments for Partial Identification of Treatment Effects

ICML 2025poster

Reliable estimation of treatment effects from observational data is important in many disciplines such as medicine. However, estimation is challenging when unconfoundedness as a standard assumption in the causal inference literature is violated. In this work, we leverage arbitrary (potentially high-…

2025

Model-agnostic meta-learners for estimating heterogeneous treatment effects over time

ICLR 2025poster

Estimating heterogeneous treatment effects (HTEs) over time is crucial in many disciplines such as personalized medicine. Existing works for this task have mostly focused on *model-based* learners that adapt specific machine-learning models and adjustment mechanisms. In contrast, model-agnostic lear…

Cited by 3SourcePDFScholar
2025

Orthogonal Survival Learners for Estimating Heterogeneous Treatment Effects from Time-to-Event Data

NeurIPS 2025poster

Estimating heterogeneous treatment effects (HTEs) is crucial for personalized decision-making. However, this task is challenging in survival analysis, which includes time-to-event data with censored outcomes (e.g., due to study dropout). In this paper, we propose a toolbox of orthogonal survival lea…

Cited by 0SourceScholar
2025

Personalized Exercise Recommendation with Semantically-Grounded Knowledge Tracing

NeurIPS 2025poster

We introduce ExRec, a general framework for personalized exercise recommendation with semantically-grounded knowledge tracing. Our method builds on the observation that existing exercise recommendation approaches simulate student performance via knowledge tracing (KT) but they often overlook two key…

Cited by 0SourcecodeScholar
2025

Stabilized Neural Prediction of Potential Outcomes in Continuous Time

ICLR 2025poster

Patient trajectories from electronic health records are widely used to estimate conditional average potential outcomes (CAPOs) of treatments over time, which then allows to personalize care. Yet, existing neural methods for this purpose have a key limitation: while some adjust for time-varying confo…

2025

Treatment Effect Estimation for Optimal Decision-Making

NeurIPS 2025poster

Decision-making across various fields, such as medicine, heavily relies on conditional average treatment effects (CATEs). Practitioners commonly make decisions by checking whether the estimated CATE is positive, even though the decision-making performance of modern CATE estimators is poorly understo…

Cited by 0SourceScholar
2024

A Neural Framework for Generalized Causal Sensitivity Analysis

ICLR 2024poster

Unobserved confounding is common in many applications, making causal inference from observational data challenging. As a remedy, causal sensitivity analysis is an important tool to draw causal conclusions under unobserved confounding with mathematical guarantees. In this paper, we propose NeuralCSA,…

2024

Bayesian Neural Controlled Differential Equations for Treatment Effect Estimation

ICLR 2024poster

Treatment effect estimation in continuous time is crucial for personalized medicine. However, existing methods for this task are limited to point estimates of the potential outcomes, whereas uncertainty estimates have been ignored. Needless to say, uncertainty quantification is crucial for reliable…

2024

Bounds on Representation-Induced Confounding Bias for Treatment Effect Estimation

ICLR 2024spotlight

State-of-the-art methods for conditional average treatment effect (CATE) estimation make widespread use of representation learning. Here, the idea is to reduce the variance of the low-sample CATE estimation by a (potentially constrained) low-dimensional representation. However, low-dimensional repre…

2024

Causal Fairness under Unobserved Confounding: A Neural Sensitivity Framework

ICLR 2024poster

Fairness for machine learning predictions is widely required in practice for legal, ethical, and societal reasons. Existing work typically focuses on settings without unobserved confounding, even though unobserved confounding can lead to severe violations of causal fairness and, thus, unfair predict…

Cited by 4SourcePDFScholar
2024

DiffPO: A causal diffusion model for learning distributions of potential outcomes

NeurIPS 2024poster

Predicting potential outcomes of interventions from observational data is crucial for decision-making in medicine, but the task is challenging due to the fundamental problem of causal inference. Existing methods are largely limited to point estimates of potential outcomes with no uncertain quantific…

Cited by 1SourcePDFScholar
2024

HQP: A Human-Annotated Dataset for Detecting Online Propaganda

ACL 2024findings

Online propaganda poses a severe threat to the integrity of societies. However, existing datasets for detecting online propaganda have a key limitation: they were annotated using weak labels that can be noisy and even incorrect. To address this limitation, our work makes the following contributions:…

2024

Meta-Learners for Partially-Identified Treatment Effects Across Multiple Environments

ICML 2024poster

Estimating the conditional average treatment effect (CATE) from observational data is relevant for many applications such as personalized medicine. Here, we focus on the widespread setting where the observational data come from multiple environments, such as different hospitals, physicians, or count…

2024

Quantifying Aleatoric Uncertainty of the Treatment Effect: A Novel Orthogonal Learner

NeurIPS 2024poster

Estimating causal quantities from observational data is crucial for understanding the safety and effectiveness of medical treatments. However, to make reliable inferences, medical practitioners require not only estimating averaged causal quantities, such as the conditional average treatment effect,…

2023

Contrastive Learning for Unsupervised Domain Adaptation of Time Series

ICLR 2023poster

Unsupervised domain adaptation (UDA) aims at learning a machine learning model using a labeled source domain that performs well on a similar yet different, unlabeled target domain. UDA is important in many applications such as medicine, where it is used to adapt risk scores across different patient…

2023

Estimating Average Causal Effects from Patient Trajectories

AAAI 2023technical

In medical practice, treatments are selected based on the expected causal effects on patient outcomes. Here, the gold standard for estimating causal effects are randomized controlled trials; however, such trials are costly and sometimes even unethical. Instead, medical practice is increasingly inter…

2023

Estimating Conditional Average Treatment Effects with Missing Treatment Information

AISTATS 2023poster

Estimating conditional average treatment effects (CATE) is challenging, especially when treatment information is missing. Although this is a widespread problem in practice, CATE estimation with missing treatments has received little attention. In this paper, we analyze CATE estimation in the setting…

2023

Estimating individual treatment effects under unobserved confounding using binary instruments

ICLR 2023poster

Estimating conditional average treatment effects (CATEs) from observational data is relevant in many fields such as personalized medicine. However, in practice, the treatment assignment is usually confounded by unobserved variables and thus introduces bias. A remedy to remove the bias is the use of…

2023

Normalizing Flows for Interventional Density Estimation

ICML 2023poster

Existing machine learning methods for causal inference usually estimate quantities expressed via the mean of potential outcomes (e.g., average treatment effect). However, such quantities do not capture the full information about the distribution of potential outcomes. In this work, we estimate the d…

2023

Partial Counterfactual Identification of Continuous Outcomes with a Curvature Sensitivity Model

NeurIPS 2023spotlight

Counterfactual inference aims to answer retrospective "what if" questions and thus belongs to the most fine-grained type of inference in Pearl's causality ladder. Existing methods for counterfactual inference with continuous outcomes aim at point identification and thus make strong and unnatural ass…

2023

Reliable Off-Policy Learning for Dosage Combinations

NeurIPS 2023poster

Decision-making in personalized medicine such as cancer therapy or critical care must often make choices for dosage combinations, i.e., multiple continuous treatments. Existing work for this task has modeled the effect of multiple treatments independently, while estimating the joint effect has recei…

2023

Sharp Bounds for Generalized Causal Sensitivity Analysis

NeurIPS 2023poster

Causal inference from observational data is crucial for many disciplines such as medicine and economics. However, sharp bounds for causal effects under relaxations of the unconfoundedness assumption (causal sensitivity analysis) are subject to ongoing research. So far, works with sharp bounds are re…

2022

Causal Transformer for Estimating Counterfactual Outcomes

ICML 2022spotlight

Estimating counterfactual outcomes over time from observational data is relevant for many applications (e.g., personalized medicine). Yet, state-of-the-art methods build upon simple long short-term memory (LSTM) networks, thus rendering inferences for complex, long-range dependencies challenging. In…

2022

Generalizing off-policy learning under sample selection bias

UAI 2022poster

Learning personalized decision policies that generalize to the target population is of great relevance. Since training data is often not representative of the target population, standard policy learning methods may yield policies that do not generalize target population. To address this challenge, w…

Cited by 32SourcePDFScholar
2022

Interpretable Off-Policy Learning via Hyperbox Search

ICML 2022spotlight

Personalized treatment decisions have become an integral part of modern medicine. Thereby, the aim is to make treatment decisions based on individual patient characteristics. Numerous methods have been developed for learning such policies from observational data that achieve the best outcome across…

2022

QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive Adaptation

EMNLP 2022main

Question answering (QA) has recently shown impressive results for answering questions from customized domains. Yet, a common challenge is to adapt QA models to an unseen target domain. In this paper, we propose a novel self-supervised framework called QADA for QA domain adaptation. QADA introduces a…

2021

Contrastive Domain Adaptation for Question Answering using Limited Text Corpora

EMNLP 2021main

Question generation has recently shown impressive results in customizing question answering (QA) systems to new domains. These approaches circumvent the need for manually annotated training data from the new domain and, instead, generate synthetic question-answer pairs that are used for training. Ho…

2021

DocParser: Hierarchical Document Structure Parsing from Renderings

AAAI 2021technical

Translating renderings (e. g. PDFs, scans) into hierarchical document structures is extensively demanded in the daily routines of many real-world applications. However, a holistic, principled approach to inferring the complete hierarchical structure in documents is missing. As a remedy, we developed…

2020

IntKB: A Verifiable Interactive Framework for Knowledge Base Completion

COLING 2020main

Knowledge bases (KBs) are essential for many downstream NLP tasks, yet their prime shortcoming is that they are often incomplete. State-of-the-art frameworks for KB completion often lack sufficient accuracy to work fully automated without human supervision. As a remedy, we propose : a novel interact…