← Search

Swetha Ganesh

6 accepted papers

2025

A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach

ICML 2025poster

This work examines average-reward reinforcement learning with general policy parametrization. Existing state-of-the-art (SOTA) guarantees for this problem are either suboptimal or hindered by several challenges, including poor scalability with respect to the size of the state-action space, high iter…

Cited by 0SourcePDFScholar
2025

Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm

NeurIPS 2025poster

This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) under general parametrized policies with smooth and bounded policy gradients. We propose a Primal-Dual Natural Actor-Critic algorithm that adeptly manages constraints while ensuring a high convergen…

Cited by 0SourceScholar
2025

Order-Optimal Global Convergence for Actor-Critic with General Policy and Neural Critic Parametrization

UAI 2025

This paper addresses the challenge of achieving order-optimal sample complexity in reinforcement learning for discounted Markov Decision Processes (MDPs) with general policy parameterization and multi-layer neural network critics. Existing approaches either fail to achieve the optimal rate or assume

2025

Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs

AISTATS 2025poster

We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring an expected regret of the order $\tilde{\mathcal{O}}(T^{2/3…

Cited by 0SourceScholar
2023

Does Momentum Help in Stochastic Optimization? A Sample Complexity Analysis.

UAI 2023poster

Stochastic Heavy Ball (SHB) and Nesterov’s Accelerated Stochastic Gradient (ASG) are popular momentum methods in optimization. While the benefits of these acceleration ideas in deterministic settings are well understood, their advantages in stochastic optimization are unclear. Several works have rec…

Cited by 4SourcePDFScholar