2025
Improved Regret and Contextual Linear Extension for Pandora's Box and Prophet Inequality
NeurIPS 2025poster
We study the Pandora’s Box problem in an online learning setting with semi-bandit feedback. In each round, the learner sequentially pays to open up to $n$ boxes with unknown reward distributions, observes rewards upon opening, and decides when to stop. The utility of the learner is the maximum obser…