← Search

P. R. Kumar

5 accepted papers

2024

Provable Policy Gradient Methods for Average-Reward Markov Potential Games

AISTATS 2024poster

We study Markov potential games under the infinite horizon average reward criterion. Most previous studies have been for discounted rewards. We prove that both algorithms based on independent policy gradient and independent natural policy gradient converge globally to a Nash equilibrium for the aver…

Cited by 8SourcePDFScholar
2021

Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits

AAAI 2021technical

Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit trade-off in linear bandits problems as well as generalized linear bandits problems. We develop novel index policies that w…

Cited by 15SourcePDFScholar
2019

Stay With Me: Lifetime Maximization Through Heteroscedastic Linear Bandits With Reneging

ICML 2019oral

Sequential decision making for lifetime maximization is a critical problem in many real-world applications, such as medical treatment and portfolio selection. In these applications, a “reneging” phenomenon, where participants may disengage from future interactions after observing an unsatisfiable ou…

Cited by 5SourcePDFScholar
2017

MT-LQG: Multi-agent planning in belief space via trajectory-optimized LQG

ICRA 2017poster

Belief space planning is concerned with the problem of finding the control policy under process and measurement uncertainties. Formulated as a stochastic control problem, the solution of a general Decentralized Partially Observed Markov Decision Process (Dec-POMDP) is a collection of feedback polici…

Cited by 3SourceScholar
2017

T-LQG: Closed-loop belief space planning via trajectory-optimized LQG

ICRA 2017poster

Planning under motion and observation uncertainties requires the solution of a stochastic control problem in the space of feedback policies. In this paper, by restricting the policy class to the linear feedback polices, we reduce the general (n2 + n)-dimensional belief space planning problem to an (…

Cited by 23SourceScholar