2024
Distributed Stochastic Contextual Bandits for Protein Drug Interaction
ICASSP 2024accepted
In recent work [1], we developed a distributed stochastic multi-arm contextual bandit algorithm to learn optimal actions when the contexts are unknown, and M agents work collaboratively under the coordination of a central server to minimize the total regret. In our model, the agents observe only the…