2026
Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract)
AAAI 2026technical
The contextual multi-armed bandit problem underlies applications in recommendations, e-commerce, finance, and healthcare, where balancing exploration and exploitation is critical. While algorithms such as Upper Confidence Bound (UCB) and Thompson Sampling (TS) achieve strong theoretical guarantees,