← Search

Mihailo Jovanovic

3 accepted papers

2022

Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic Convergence

ICML 2022oral

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPGs). To learn a Nash equilibrium of an MPG in which the size of state space and/or the number of players can be very large, we propose…

Cited by 97SourcePDFScholar
2021

Provably Efficient Safe Exploration via Primal-Dual Policy Optimization

AISTATS 2021poster

We study the safe reinforcement learning problem using the constrained Markov decision processes in which an agent aims to maximize the expected total reward subject to a safety constraint on the expected total value of a utility function. We focus on an episodic setting with the function approximat…

Cited by 200SourcePDFScholar
2020

Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision Processes

NeurIPS 2020poster

We study sequential decision-making problems in which each agent aims to maximize the expected total reward while satisfying a constraint on the expected total utility. We employ the natural policy gradient method to solve the discounted infinite-horizon Constrained Markov Decision Processes (CMDPs)…

Cited by 241SourcePDFScholar