← Search

Sanath Kumar Krishnamurthy

6 accepted papers

2025

Selective Uncertainty Propagation in Offline RL

AAAI 2025technical

We consider the finite-horizon offline reinforcement learning (RL) setting, and are motivated by the challenge of learning the policy at any step h in dynamic programming (DP) algorithms. To learn this, it is sufficient to evaluate the treatment effect of deviating from the behavioral policy at step…

Cited by 1SourcePDFScholar
2024

Towards Costless Model Selection in Contextual Bandits: A Bias-Variance Perspective

AISTATS 2024poster

Model selection in supervised learning provides costless guarantees as if the model that best balances bias and variance was known a priori. We study the feasibility of similar guarantees for cumulative regret minimization in the stochastic contextual bandit setting. Recent work [Marinov and Zimmert…

Cited by 3SourcePDFScholar
2023

Flexible and Efficient Contextual Bandits with Heterogeneous Treatment Effect Oracles

AISTATS 2023poster

Contextual bandit algorithms often estimate reward models to inform decision-making. However, true rewards can contain action-independent redundancies that are not relevant for decision-making. We show it is more data-efficient to estimate any function that explains the reward differences between ac…

2023

Proportional Response: Contextual Bandits for Simple and Cumulative Regret Minimization

NeurIPS 2023poster

In many applications, e.g. in healthcare and e-commerce, the goal of a contextual bandit may be to learn an optimal treatment assignment policy at the end of the experiment. That is, to minimize simple regret. However, this objective remains understudied. We propose a new family of computationally e…

Cited by 14SourcePDFScholar
2021

Adapting to misspecification in contextual bandits with offline regression oracles

ICML 2021spotlight

Computationally efficient contextual bandits are often based on estimating a predictive model of rewards given contexts and arms using past data. However, when the reward model is not well-specified, the bandit algorithm may incur unexpected regret, so recent work has focused on algorithms that are…

Cited by 28SourcePDFScholar