← Search

Aurelien Bibaut

7 accepted papers

2024

Inferring the Long-Term Causal Effects of Long-Term Treatments from Short-Term Experiments

ICML 2024oral

We study inference on the long-term causal effect of a continual exposure to a novel intervention, which we term a long-term treatment, based on an experiment involving only short-term observations. Key examples include the long-term health effects of regularly-taken medicine or of environmental haz…

2021

Post-Contextual-Bandit Inference

NeurIPS 2021poster

Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies. To support credible inference on novel intervent…

Cited by 54SourcePDFScholar
2021

Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy Learning

NeurIPS 2021poster

Empirical risk minimization (ERM) is the workhorse of machine learning, whether for classification and regression or for off-policy policy learning, but its model-agnostic guarantees can fail when we use adaptively collected data, such as the result of running a contextual bandit algorithm. We study…

Cited by 17SourcePDFScholar
2019

More Efficient Off-Policy Evaluation through Regularized Targeted Learning

ICML 2019oral

We study the problem of off-policy evaluation (OPE) in Reinforcement Learning (RL), where the aim is to estimate the performance of a new policy given historical data that may have been generated by a different policy, or policies. In particular, we introduce a novel doubly-robust estimator for the…

Cited by 42SourcePDFScholar
2019

On the Design of Estimators for Bandit Off-Policy Evaluation

ICML 2019oral

Off-policy evaluation is the problem of estimating the value of a target policy using data collected under a different policy. Given a base estimator for bandit off-policy evaluation and a parametrized class of control variates, we address the problem of computing a control variate in that class tha…

Cited by 36SourcePDFScholar