2022
An explore-then-commit algorithm for submodular maximization under full-bandit feedback
UAI 2022poster
We investigate the problem of combinatorial multi-armed bandits with stochastic submodular (in expectation) rewards and full-bandit feedback, where no extra information other than the reward of selected action at each time step $t$ is observed. We propose a simple algorithm, Explore-Then-Commit Gree…