← Search

Debajoy Mukherjee

1 accepted papers

2025

DOPL: Direct Online Preference Learning for Restless Bandits with Preference Feedback

ICLR 2025poster

Restless multi-armed bandits (RMAB) has been widely used to model constrained sequential decision making problems, where the state of each restless arm evolves according to a Markov chain and each state transition generates a scalar reward. However, the success of RMAB crucially relies on the availa…

Cited by 2SourcePDFScholar