2024
Learning the Pareto Set Under Incomplete Preferences: Pure Exploration in Vector Bandits
AISTATS 2024poster
We study pure exploration in bandit problems with vector-valued rewards, where the goal is to (approximately) identify the Pareto set of arms given incomplete preferences induced by a polyhedral convex cone. We address the open problem of designing sample-efficient learning algorithms for such probl…