← Search

Gilles Stoltz

6 accepted papers

2025

Narrowing the Gap between Adversarial and Stochastic MDPs via Policy Optimization

AISTATS 2025poster

We consider the problem of learning in adversarial Markov decision processes [MDPs] with an oblivious adversary in a full-information setting. The agent interacts with an environment during $T$ episodes, each of which consists of $H$ stages, and each episode is evaluated with respect to a reward fun…

Cited by 0SourceScholar
2023

Small Total-Cost Constraints in Contextual Bandits with Knapsacks, with Application to Fairness

NeurIPS 2023poster

We consider contextual bandit problems with knapsacks [CBwK], a problem where at each round, a scalar reward is obtained and vector-valued costs are suffered. The learner aims to maximize the cumulative rewards while ensuring that the cumulative costs are lower than some predetermined cost constrain…

Cited by 2SourcePDFScholar
2021

A Unified Approach to Fair Online Learning via Blackwell Approachability

NeurIPS 2021spotlight

We provide a setting and a general approach to fair online learning with stochastic sensitive and non-sensitive contexts. The setting is a repeated game between the Player and Nature, where at each stage both pick actions based on the contexts. Inspired by the notion of unawareness, we assume that t…

Cited by 11SourcePDFScholar
2019

Target Tracking for Contextual Bandits: Application to Demand Side Management

ICML 2019oral

We propose a contextual-bandit approach for demand side management by offering price incentives. More precisely, a target mean consumption is set at each round and the mean consumption is modeled as a complex function of the distribution of prices sent and of some contextual variables such as the te…

Cited by 14SourcePDFScholar