← Search

Maria Dimakopoulou

9 accepted papers

2025

Concurrent Reinforcement Learning with Aggregated States via Randomized Least Squares Value Iteration

ICML 2025poster

Designing learning agents that explore efficiently in a complex environment has been widely recognized as a fundamental challenge in reinforcement learning. While a number of works have demonstrated the effectiveness of techniques based on randomized value functions on a single agent, it remains un…

Cited by 0SourcePDFScholar
2022

Society of Agents: Regret Bounds of Concurrent Thompson Sampling

NeurIPS 2022accept

We consider the concurrent reinforcement learning problem where $n$ agents simultaneously learn to make decisions in the same environment by sharing experience with each other. Existing works in this emerging area have empirically demonstrated that Thompson sampling (TS) based algorithms provide a…

Cited by 5SourcePDFScholar
2021

Post-Contextual-Bandit Inference

NeurIPS 2021poster

Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies. To support credible inference on novel intervent…

Cited by 54SourcePDFScholar
2021

Risk Minimization from Adaptively Collected Data: Guarantees for Supervised and Policy Learning

NeurIPS 2021poster

Empirical risk minimization (ERM) is the workhorse of machine learning, whether for classification and regression or for off-policy policy learning, but its model-agnostic guarantees can fail when we use adaptively collected data, such as the result of running a contextual bandit algorithm. We study…

Cited by 17SourcePDFScholar
2020

Doubly robust off-policy evaluation with shrinkage

ICML 2020poster

We propose a new framework for designing estimators for off-policy evaluation in contextual bandits. Our approach is based on the asymptotically optimal doubly robust estimator, but we shrink the importance weights to minimize a bound on the mean squared error, which results in a better bias-varianc…

Cited by 119SourcePDFScholar
2019

On the Design of Estimators for Bandit Off-Policy Evaluation

ICML 2019oral

Off-policy evaluation is the problem of estimating the value of a target policy using data collected under a different policy. Given a base estimator for bandit off-policy evaluation and a parametrized class of control variates, we address the problem of computing a control variate in that class tha…

Cited by 36SourcePDFScholar
2018

Scalable Coordinated Exploration in Concurrent Reinforcement Learning

NeurIPS 2018poster

We consider a team of reinforcement learning agents that concurrently operate in a common environment, and we develop an approach to efficient coordinated exploration that is suitable for problems of practical scale. Our approach builds on the seed sampling concept introduced in Dimakopoulou and Van…

Cited by 61SourcePDFScholar