← Search

Matthieu Martin

5 accepted papers

2023

Sequential Counterfactual Risk Minimization

ICML 2023poster

Counterfactual Risk Minimization (CRM) is a framework for dealing with the logged bandit feedback problem, where the goal is to improve a logging policy using offline data. In this paper, we explore the case where it is possible to deploy learned policies multiple times and acquire new data. We exte…

2022

Efficient Kernelized UCB for Contextual Bandits

AISTATS 2022poster

In this paper, we tackle the computational efficiency of kernelized UCB algorithms in contextual bandits. While standard methods require a $\mathcal{O}(CT^3)$ complexity where $T$ is the horizon and the constant $C$ is related to optimizing the UCB rule, we propose an efficient contextual algorithm…

Cited by 24SourcePDFScholar
2021

Zeroth-Order Non-Convex Learning via Hierarchical Dual Averaging

ICML 2021spotlight

We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization {–} i.e., learning processes where, at each stage, the optimizer is facing an unknown non-convex loss function and only receives the incurred loss as feedback. The proposed class of policies relies on…

Cited by 16SourcePDFScholar
2020

Online Non-Convex Optimization with Imperfect Feedback

NeurIPS 2020poster

We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes – or otherwise constructs – an inexact model for the loss function encountered at each stage, and we propose a mixed-strategy learning policy based on dual averaging. In this…

Cited by 26SourcePDFScholar