GINO-Q: Learning an Asymptotically Optimal Index Policy for Restless Multi-armed Bandits
The restless multi-armed bandit (RMAB) framework is a popular model with applications across a wide variety of fields. However, its solution is hindered by the exponentially growing state space (with respect to the number of arms) and the combinatorial action space, making traditional reinforcement