← Search

Yashaswini Murthy

2 accepted papers

2025

Global Convergence of Policy Gradient in Average Reward MDPs

ICLR 2025poster

We present the first comprehensive finite-time global convergence analysis of policy gradient for infinite horizon average reward Markov decision processes (MDPs). Specifically, we focus on ergodic tabular MDPs with finite state and action spaces. Our analysis shows that the policy gradient iterates…

Cited by 0SourcePDFScholar
2023

Performance Bounds for Policy-Based Average Reward Reinforcement Learning Algorithms

NeurIPS 2023poster

Many policy-based reinforcement learning (RL) algorithms can be viewed as instantiations of approximate policy iteration (PI), i.e., where policy improvement and policy evaluation are both performed approximately. In applications where the average reward objective is the meaningful performance metri…

Cited by 4SourcePDFScholar