← Search

Mudit Gaur

3 accepted papers

2025

On the Sample Complexity Bounds of Bilevel Reinforcement Learning

NeurIPS 2025poster

Bilevel reinforcement learning (BRL) has emerged as a powerful framework for aligning generative models, yet its theoretical foundations, especially sample complexity bounds, remain underexplored. In this work, we present the first sample complexity bound for BRL, establishing a rate of $\mathcal{O}…

Cited by 0SourceScholar
2024

Closing the Gap: Achieving Global Convergence (Last Iterate) of Actor-Critic under Markovian Sampling with Neural Network Parametrization

ICML 2024spotlight

The current state-of-the-art theoretical analysis of Actor-Critic (AC) algorithms significantly lags in addressing the practical aspects of AC implementations. This crucial gap needs bridging to bring the analysis in line with practical implementations of AC. To address this, we advocate for conside…

Cited by 2SourcePDFScholar
2023

On the Global Convergence of Fitted Q-Iteration with Two-layer Neural Network Parametrization

ICML 2023poster

Deep Q-learning based algorithms have been applied successfully in many decision making problems, while their theoretical foundations are not as well understood. In this paper, we study a Fitted Q-Iteration with two-layer ReLU neural network parameterization, and find the sample complexity guarantee…

Cited by 3SourcePDFScholar