← Search

Wang Chi Cheung

13 accepted papers

2026

Efficiently Solving Discounted MDPs via Predictions with Unknown Prediction Errors

ICML 2026poster

We study infinite-horizon discounted Markov decision processes (DMDPs) under a generative model. Motivated by the Algorithms with Advice framework (Mitzenmacher and Vassilvitskii, 2022), we propose a novel framework to investigate how black-box predictions of the transition matrix can enhance sample…

Cited by 0SourceScholar
2021

Probabilistic Sequential Shrinking: A Best Arm Identification Algorithm for Stochastic Bandits with Corruptions

ICML 2021spotlight

We consider a best arm identification (BAI) problem for stochastic bandits with adversarial corruptions in the fixed-budget setting of T steps. We design a novel randomized algorithm, Probabilistic Sequential Shrinking(u) (PSS(u)), which is agnostic to the amount of corruptions. When the amount of c…

2020

Best Arm Identification for Cascading Bandits in the Fixed Confidence Setting

ICML 2020poster

We design and analyze CascadeBAI, an algorithm for finding the best set of K items, also called an arm, within the framework of cascading bandits. An upper bound on the time complexity of CascadeBAI is derived by overcoming a crucial analytical challenge, namely, that of probabilistically estimating…

Cited by 11SourcePDFScholar
2020

Reinforcement Learning for Non-Stationary Markov Decision Processes: The Blessing of (More) Optimism

ICML 2020poster

We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, \ie, both the reward and state transition distributions are allowed to evolve over time, as long as their respective total variations, quantified by suitable metrics, do not exc…

Cited by 128SourcePDFScholar
2019

Regret Minimization for Reinforcement Learning with Vectorial Feedback and Complex Objectives

NeurIPS 2019poster

We consider an agent who is involved in an online Markov decision process, and receives a vector of outcomes every round. The agent aims to simultaneously optimize multiple objectives associated with the multi-dimensional outcomes. Due to state transitions, it is challenging to balance the vectoria…