← Search

Yu-Heng Hung

2 accepted papers

2023

Reward-Biased Maximum Likelihood Estimation for Neural Contextual Bandits: A Distributional Learning Perspective

AAAI 2023technical

Reward-biased maximum likelihood estimation (RBMLE) is a classic principle in the adaptive control literature for tackling explore-exploit trade-offs. This paper studies the neural contextual bandit problem from a distributional perspective and proposes NeuralRBMLE, which leverages the likelihood of…

Cited by 2SourcePDFScholar
2021

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

AAAI 2021technical

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that w…

Cited by 15SourcePDFScholar