2026
Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
ICML 2026spotlight
We study linear contextual bandits under rare parameter updates: the learner may incorporate reward feedback into its parameter estimate only at a small number of update times, while still observing contexts online and selecting actions sequentially. This viewpoint clarifies a practical distinction …