← Search

Emile Timothy Anand

4 accepted papers

2025

Feel-Good Thompson Sampling for Contextual Bandits: a Markov Chain Monte Carlo Showdown

NeurIPS 2025poster

Thompson Sampling (TS) is widely used to address the exploration/exploitation tradeoff in contextual bandits, yet recent theory shows that it does not explore aggressively enough in high-dimensional problems. Feel-Good Thompson Sampling (FG-TS) addresses this by adding an optimism bonus that biases…

Cited by 0SourcecodeScholar
2025

Mean-Field Sampling for Cooperative Multi-Agent Reinforcement Learning

NeurIPS 2025spotlight

Designing efficient algorithms for multi-agent reinforcement learning (MARL) is fundamentally challenging because the size of the joint state and action spaces grows exponentially in the number of agents. These difficulties are exacerbated when balancing sequential global decision-making with local…

Cited by 0SourceScholar
2023

Online Adaptive Policy Selection in Time-Varying Systems: No-Regret via Contractive Perturbations

NeurIPS 2023poster

We study online adaptive policy selection in systems with time-varying costs and dynamics. We develop the Gradient-based Adaptive Policy Selection (GAPS) algorithm together with a general analytical framework for online policy selection via online optimization. Under our proposed notion of contracti…

Cited by 16SourcePDFScholar