← Search

Siva Theja Maguluri

8 accepted papers

2025

Stochastic Approximation with Unbounded Markovian Noise: A General-Purpose Theorem

AISTATS 2025poster

Motivated by engineering applications such as resource allocation in networks and inventory systems, we consider average-reward Reinforcement Learning with unbounded state space and reward function. Recent work Murthy et al. 2024 studied this problem in the actor-critic framework and established fin…

Cited by 0SourceScholar
2024

Policy Evaluation for Variance in Average Reward Reinforcement Learning

ICML 2024poster

We consider an average reward reinforcement learning (RL) problem and work with asymptotic variance as a risk measure to model safety-critical applications. We design a temporal-difference (TD) type algorithm tailored for policy evaluation in this context. Our algorithm is based on linear stochastic…

Cited by 5SourcePDFScholar
2022

Federated Reinforcement Learning: Linear Speedup Under Markovian Sampling

ICML 2022oral

Since reinforcement learning algorithms are notoriously data-intensive, the task of sampling observations from the environment is usually split across multiple agents. However, transferring these observations from the agents to a central location can be prohibitively expensive in terms of the commun…

Cited by 87SourcePDFScholar
2022

Sample Complexity of Policy-Based Methods under Off-Policy Sampling and Linear Function Approximation

AISTATS 2022poster

In this work, we study policy-based methods for solving the reinforcement learning problem, where off-policy sampling and linear function approximation are employed for policy evaluation, and various policy update rules (including natural policy gradient) are considered for policy improvement. To so…

Cited by 23SourcePDFScholar
2021

Finite Sample Analysis of Average-Reward TD Learning and $Q$-Learning

NeurIPS 2021poster

The focus of this paper is on sample complexity guarantees of average-reward reinforcement learning algorithms, which are known to be more challenging to study than their discounted-reward counterparts. To the best of our knowledge, we provide the first known finite sample guarantees using both cons…

Cited by 36SourcePDFScholar
2021

Finite-Sample Analysis of Off-Policy Natural Actor-Critic Algorithm

ICML 2021spotlight

In this paper, we provide finite-sample convergence guarantees for an off-policy variant of the natural actor-critic (NAC) algorithm based on Importance Sampling. In particular, we show that the algorithm converges to a global optimal policy with a sample complexity of $\mathcal{O}(\epsilon^{-3}\log…

Cited by 34SourcePDFScholar
2021

Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators

NeurIPS 2021poster

In TD-learning, off-policy sampling is known to be more practical than on-policy sampling, and by decoupling learning from data collection, it enables data reuse. It is known that policy evaluation has the interpretation of solving a generalized Bellman equation. In this paper, we derive finite-samp…

Cited by 16SourcePDFScholar
2020

Finite-Sample Analysis of Contractive Stochastic Approximation Using Smooth Convex Envelopes

NeurIPS 2020poster

Stochastic Approximation (SA) is a popular approach for solving fixed-point equations where the information is corrupted by noise. In this paper, we consider an SA involving a contraction mapping with respect to an arbitrary norm, and show its finite-sample error bounds while using different stepsiz…

Cited by 66SourcePDFScholar