2025
Neural Combinatorial Clustered Bandits for Recommendation Systems
AAAI 2025technical
We consider the contextual combinatorial bandit setting where in each round, the learning agent, e.g., a recommender system, selects a subset of "arms,'' e.g., products, and observes rewards for both the individual base arms, which are a function of known features (called "context''), and the super…