2020
Multi-Feedback Bandit Learning with Probabilistic Contexts
IJCAI 2020poster
Contextual bandit is a classic multi-armed bandit setting, where side information (i.e., context) is available before arm selection. A standard assumption is that exact contexts are perfectly known prior to arm selection and only single feedback is returned. In this work, we focus on multi-feedback…