Linear Bandits beyond Inner Product Spaces, the case of Bandit Optimal Transport
Linear bandits have long been a central topic in online learning, with applications ranging from recommendation systems to adaptive clinical trials. Their general learnability has been established when the objective is to minimise the inner product between a cost parameter and the decision variable.…