2026
Provably Efficient Multi-Objective Bandit Algorithms Under Preference-Centric Customization
AAAI 2026technical
Multi-objective multi-armed bandit (MO-MAB) problems traditionally aim to achieve Pareto optimality. However, real-world scenarios often involve users with varying preferences across objectives, resulting in a Pareto-optimal arm that may score high for one user but perform quite poorly for another.