Beyond the Lower Bound: Bridging Regret Minimization and Best Arm Identification in Lexicographic Bandits
In multi-objective decision-making with hierarchical preferences, lexicographic bandits provide a natural framework for optimizing multiple objectives in a prioritized order. In this setting, a learner repeatedly selects arms and observes reward vectors, aiming to maximize the reward for the highest