Learning the Pareto Front Using Bootstrapped Observation Samples
We consider Pareto front identification~(PFI) for linear bandits (PFILin), i.e., the goal is to identify a set of arms with undominated mean reward vectors when the mean reward vector is a linear function of the context. PFILin includes the best arm identification problem and multi-objective activ…