← Search

Prashanth L. A.

4 accepted papers

2025

Risk-sensitive Bandits: Arm Mixture Optimality and Regret-efficient Algorithms

AISTATS 2025poster

This paper introduces a general framework for risk-sensitive bandits that integrates the notions of risk-sensitive objectives by adopting a rich class of {\em distortion riskmetrics}. The introduced framework subsumes the various existing risk-sensitive models. An important and hitherto unknown obse…

Cited by 0SourcecodeScholar
2024

Policy Evaluation for Variance in Average Reward Reinforcement Learning

ICML 2024poster

We consider an average reward reinforcement learning (RL) problem and work with asymptotic variance as a risk measure to model safety-critical applications. We design a temporal-difference (TD) type algorithm tailored for policy evaluation in this context. Our algorithm is based on linear stochastic…

Cited by 5SourcePDFScholar