← Search

Shubhada Agrawal

5 accepted papers

2026

Asymptotically Optimal Sequential Testing with Markovian Data

ICML 2026poster

We study one-sided and $\alpha$-correct sequential hypothesis testing for data generated by an ergodic Markov chain. The *null* hypothesis is that the unknown transition matrix belongs to a prescribed set $\cal P$ of stochastic matrices, and the *alternative* corresponds to a disjoint set $\cal Q$. …

Cited by 1SourceScholar
2024

Optimal Top-Two Method for Best Arm Identification and Fluid Analysis

NeurIPS 2024poster

Top-2 methods have become popular in solving the best arm identification (BAI) problem. The best arm, or the arm with the largest mean amongst finitely many, is identified through an algorithm that at any sequential step independently pulls the empirical best arm, with a fixed probability $\beta$, a…

Cited by 3SourcePDFScholar
2024

Policy Evaluation for Variance in Average Reward Reinforcement Learning

ICML 2024poster

We consider an average reward reinforcement learning (RL) problem and work with asymptotic variance as a risk measure to model safety-critical applications. We design a temporal-difference (TD) type algorithm tailored for policy evaluation in this context. Our algorithm is based on linear stochastic…

Cited by 5SourcePDFScholar
2021

Optimal Best-Arm Identification Methods for Tail-Risk Measures

NeurIPS 2021poster

Conditional value-at-risk (CVaR) and value-at-risk (VaR) are popular tail-risk measures in finance and insurance industries as well as in highly reliable, safety-critical uncertain environments where often the underlying probability distributions are heavy-tailed. We use the multi-armed bandit best-…

Cited by 53SourcePDFScholar