IJCAI 20260 citations

Safe Multi-Objective Linear Bandits with Hierarchical Preferences

Bo Xue, Mengxia He, Yilu Liu, Ji Cheng, Zhe Zhao, Qingfu Zhang

Abstract

Multi-objective bandits with hierarchical preferences and safety constraints is central to many real-world decision-making tasks such as healthcare treatment planning and safe autonomous control, where multiple objectives must be optimized according to their priorities while ensuring safety requirements are satisfied. In this paper, we study a multi-objective stochastic linear bandit framework that incorporates hierarchical preferences together with safety constraints, requiring the learner to remain competitive with respect to a known baseline policy. We consider two practically motivated safety models: (i) cumulative constraints, which require the cumulative performance to exceed the baseline, and (ii) stage-wise constraints, which impose this requirement at each time step. We propose two algorithms, LexUCB-C and LexTS-S, designed for the cumulative and stage-wise settings, respectively. We establish regret bounds showing that both algorithms achieve performance comparable to existing single-objective safe linear bandit methods, while simultaneously optimizing multiple objectives. In addition to theoretical guarantees, we develop a carefully designed experimental framework that captures the interaction between hierarchical preferences and safety constraints. Experiments on synthetic and real-world datasets validate our theory and demonstrate the effectiveness of the proposed methods.

Constraint Satisfaction and Optimization: Constraint optimization problemsAI: Uncertainty in AIMachine Learning: Learning theory
BibTeX
@inproceedings{ijcai2026_safemultiobjecti,
  title = {Safe Multi-Objective Linear Bandits with Hierarchical Preferences},
  author = {Bo Xue and Mengxia He and Yilu Liu and Ji Cheng and Zhe Zhao and Qingfu Zhang},
  booktitle = {IJCAI 2026},
  year = {2026}
}
Safe Multi-Objective Linear Bandits with Hierarchical Preferences · IJCAI 2026