← Search

Shipra Agrawal

8 accepted papers

2018

Bandits with Delayed, Aggregated Anonymous Feedback

ICML 2018oral

We study a variant of the stochastic $K$-armed bandit problem, which we call "bandits with delayed, aggregated anonymous feedback”. In this problem, when the player pulls an arm, a reward is generated, however it is not immediately observed. Instead, at the end of each round the player observes only…

Cited by 148SourcePDFScholar
2018

Proportional Allocation: Simple, Distributed, and Diverse Matching with High Entropy

ICML 2018oral

Inspired by many applications of bipartite matching in online advertising and machine learning, we study a simple and natural iterative proportional allocation algorithm: Maintain a priority score $\priority_a$ for each node $a\in \mathds{A}$ on one side of the bipartition, initialized as $\priority…

Cited by 41SourcePDFScholar
2017

Optimistic posterior sampling for reinforcement learning: worst-case regret bounds

NeurIPS 2017poster

We present an algorithm based on posterior sampling (aka Thompson sampling) that achieves near-optimal worst-case regret bounds when the underlying Markov Decision Process (MDP) is communicating with a finite, though unknown, diameter. Our main result is a high probability regret upper bound of $\ti…

Cited by 267SourcePDFScholar