← Search

Xuedong Shang

3 accepted papers

2021

UCB Momentum Q-learning: Correcting the bias without forgetting

ICML 2021oral

We propose UCBMQ, Upper Confidence Bound Momentum Q-learning, a new algorithm for reinforcement learning in tabular and possibly stage-dependent, episodic Markov decision process. UCBMQ is based on Q-learning where we add a momentum term and rely on the principle of optimism in face of uncertainty t…

2020

Fixed-confidence guarantees for Bayesian best-arm identification

AISTATS 2020poster

We investigate and provide new insights on the sampling rule called Top-Two Thompson Sampling (TTTS). In particular, we justify its use for fixed-confidence best-arm identification. We further propose a variant of TTTS called Top-Two Transportation Cost (T3C), which disposes of the computational bur…

Cited by 83SourcePDFScholar