← Search

Washim Uddin Mondal

9 accepted papers

2025

A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic Approach

ICML 2025poster

This work examines average-reward reinforcement learning with general policy parametrization. Existing state-of-the-art (SOTA) guarantees for this problem are either suboptimal or hindered by several challenges, including poor scalability with respect to the size of the state-action space, high iter…

Cited by 0SourcePDFScholar
2025

Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning

NeurIPS 2025poster

We present the first finite-sample analysis of policy evaluation in robust average-reward Markov Decision Processes (MDPs). Prior work in this setting have established only asymptotic convergence guarantees, leaving open the question of sample complexity. In this work, we address this gap by showing…

Cited by 0SourceScholar
2025

Global Convergence for Average Reward Constrained MDPs with Primal-Dual Actor Critic Algorithm

NeurIPS 2025poster

This paper investigates infinite-horizon average reward Constrained Markov Decision Processes (CMDPs) under general parametrized policies with smooth and bounded policy gradients. We propose a Primal-Dual Natural Actor-Critic algorithm that adeptly manages constraints while ensuring a high convergen…

Cited by 0SourceScholar
2025

Order-Optimal Global Convergence for Actor-Critic with General Policy and Neural Critic Parametrization

UAI 2025

This paper addresses the challenge of achieving order-optimal sample complexity in reinforcement learning for discounted Markov Decision Processes (MDPs) with general policy parameterization and multi-layer neural network critics. Existing approaches either fail to achieve the optimal rate or assume

2025

Order-Optimal Regret with Novel Policy Gradient Approaches in Infinite-Horizon Average Reward MDPs

AISTATS 2025poster

We present two Policy Gradient-based algorithms with general parametrization in the context of infinite-horizon average reward Markov Decision Process (MDP). The first one employs Implicit Gradient Transport for variance reduction, ensuring an expected regret of the order $\tilde{\mathcal{O}}(T^{2/3…

Cited by 0SourceScholar
2024

Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient Algorithm

NeurIPS 2024poster

This paper explores the realm of infinite horizon average reward Constrained Markov Decision Processes (CMDPs). To the best of our knowledge, this work is the first to delve into the regret and constraint violation analysis of average reward CMDPs with a general policy parametrization. To address th…

Cited by 2SourcePDFScholar
2024

Regret Analysis of Policy Gradient Algorithm for Infinite Horizon Average Reward Markov Decision Processes

AAAI 2024technical

In this paper, we consider an infinite horizon average reward Markov Decision Process (MDP). Distinguishing itself from existing works within this context, our approach harnesses the power of the general policy gradient-based algorithm, liberating it from the constraints of assuming a linear MDP str…

Cited by 17SourcePDFScholar
2024

Sample-Efficient Constrained Reinforcement Learning with General Parameterization

NeurIPS 2024poster

We consider a constrained Markov Decision Problem (CMDP) where the goal of an agent is to maximize the expected discounted sum of rewards over an infinite horizon while ensuring that the expected discounted sum of costs exceeds a certain threshold. Building on the idea of momentum-based acceleration…

Cited by 6SourcePDFScholar
2022

Can mean field control (mfc) approximate cooperative multi agent reinforcement learning (marl) with non-uniform interaction?

UAI 2022poster

Mean-Field Control (MFC) is a powerful tool to solve Multi-Agent Reinforcement Learning (MARL) problems. Recent studies have shown that MFC can well-approximate MARL when the population size is large and the agents are exchangeable. Unfortunately, the presumption of exchangeability implies that all…