Mitigating Disparity while Maximizing Reward: Tight Anytime Guarantee for Improving Bandits
We study the Improving Multi-Armed Bandit problem, where the reward obtained from an arm increases with the number of pulls it receives. This model provides an elegant abstraction for many real-world problems in domains such as education and employment, where decisions about the distribution of oppo…