2025
Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost Subsidy
ICLR 2025poster
Multi-armed bandits (MAB) are commonly used in sequential online decision-making when the reward of each decision is an unknown random variable. In practice however, the typical goal of maximizing total reward may be less important than minimizing the total cost of the decisions taken, subject to a…