← Search

Dastan Omirzak

1 accepted papers

2026

Efficient Contextual Bandit Learning via Reward-Space Sampling and Online Optimization (Student Abstract)

AAAI 2026technical

The contextual multi-armed bandit problem underlies applications in recommendations, e-commerce, finance, and healthcare, where balancing exploration and exploitation is critical. While algorithms such as Upper Confidence Bound (UCB) and Thompson Sampling (TS) achieve strong theoretical guarantees,

Cited by 0SourcePDFScholar