← Search

Erwan Scornet

15 accepted papers

2025

A Unified Framework for the Transportability of Population-Level Causal Measures

NeurIPS 2025poster

Generalization methods offer a powerful solution to one of the key drawbacks of randomized controlled trials (RCTs): their limited representativeness. By enabling the transport of treatment effect estimates to target populations subject to distributional shifts, these methods are increasingly recogn…

Cited by 0SourceScholar
2025

Quantifying Treatment Effects: Estimating Risk Ratios via Observational Studies

ICML 2025poster

The Risk Difference (RD), an absolute measure of effect, is widely used and well-studied in both randomized controlled trials (RCTs) and observational studies. Complementary to the RD, the Risk Ratio (RR), as a relative measure, is critical for a comprehensive understanding of intervention effects:…

Cited by 0SourcePDFScholar
2024

Random features models: a way to study the success of naive imputation

ICML 2024poster

Constant (naive) imputation is still widely used in practice as this is a first easy-to-use technique to deal with missing data. Yet, this simple method could be expected to induce a large bias for prediction purposes, as the imputed input may strongly differ from the true underlying data. However,…

Cited by 6SourcePDFScholar
2023

Naive imputation implicitly regularizes high-dimensional linear models

ICML 2023poster

Two different approaches exist to handle missing values for prediction: either imputation, prior to fitting any predictive algorithms, or dedicated methods able to natively incorporate missing values. While imputation is widely (and easily) use, it is unfortunately biased when low-capacity predictor…

Cited by 9SourcePDFScholar
2022

Near-optimal rate of consistency for linear models with missing values

ICML 2022spotlight

Missing values arise in most real-world data sets due to the aggregation of multiple sources and intrinsically missing information (sensor failure, unanswered questions in surveys...). In fact, the very nature of missing values usually prevents us from running standard learning algorithms. In this p…

Cited by 12SourcePDFScholar
2022

SHAFF: Fast and consistent SHApley eFfect estimates via random Forests

AISTATS 2022poster

Interpretability of learning algorithms is crucial for applications involving critical decisions, and variable importance is one of the main interpretation tools. Shapley effects are now widely used to interpret both tree ensembles and neural networks, as they can efficiently handle dependence and i…

2021

Interpretable Random Forests via Rule Extraction

AISTATS 2021poster

We introduce SIRUS (Stable and Interpretable RUle Set) for regression, a stable rule learning algorithm, which takes the form of a short and simple list of rules. State-of-the-art learning algorithms are often referred to as “black boxes” because of the high number of operations involved in their pr…

Cited by 95SourcePDFScholar
2021

What’s a good imputation to predict with missing values?

NeurIPS 2021spotlight

How to learn a good predictor on data with missing values? Most efforts focus on first imputing as well as possible and second learning on the completed data to predict the outcome. Yet, this widespread practice has no theoretical grounding. Here we show that for almost all imputation functions, an…

2020

Linear predictor on linearly-generated data with missing values: non consistency and solutions

AISTATS 2020poster

We consider building predictors when the data have missing values. We study the seemingly-simple case where the target to predict is a linear function of the fully observed data and we show that, in the presence of missing values, the optimal predictor is not linear in general. In the particular Gau…

2020

NeuMiss networks: differentiable programming for supervised learning with missing values.

NeurIPS 2020oral

The presence of missing values makes supervised learning much more challenging. Indeed, previous work has shown that even when the response is a linear function of the complete data, the optimal predictor is a complex function of the observed entries and the missingness indicator. As a result, the c…

2017

Universal consistency and minimax rates for online Mondrian Forests

NeurIPS 2017poster

We establish the consistency of an algorithm of Mondrian Forests~\cite{lakshminarayanan2014mondrianforests,lakshminarayanan2016mondrianuncertainty}, a randomized classification algorithm that can be implemented online. First, we amend the original Mondrian Forest algorithm proposed in~\cite{lakshmi…

Cited by 25SourcePDFScholar