2025
DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback
ICLR 2025poster
Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availa…