← Search

Shalabh Bhatnagar

10 accepted papers

2025

One Encoder to Rule them All: Representation Learning for Model-free Visual Reinforcement Learning using Fourier Neural Operators

ICCV 2025poster

Representation learning lies at the core of deep reinforcement learning. Although CNNs have traditionally served as the primary models for encoding image observations, modifying the encoder architecture introduces challenges, especially due to the necessity of determining a new set of hyperparameter…

2025

Two-Timescale Critic-Actor for Average Reward MDPs with Function Approximation

AAAI 2025technical

Several recent works have focused on carrying out non-asymptotic convergence analyses for AC algorithms. Recently, a two-timescale critic-actor algorithm has been presented for the discounted cost setting in the look-up table case where the timescales of the actor and the critic are reversed and onl…

2024

A Cubic-regularized Policy Newton Algorithm for Reinforcement Learning

AISTATS 2024poster

We consider the problem of control in the setting of reinforcement learning (RL), where model information is not available. Policy gradient algorithms are a popular solution approach for this problem and are usually shown to converge to a stationary point of the value function. In this paper, we pro…

Cited by 3SourcePDFScholar
2024

Finite-Time Analysis of Three-Timescale Constrained Actor-Critic and Constrained Natural Actor-Critic Algorithms.

UAI 2024poster

Actor Critic methods have found immense applications on a wide range of Reinforcement Learning tasks especially when the state-action space is large. In this paper, we consider actor critic and natural actor critic algorithms with function approximation for constrained Markov decision processes (C-M…

2023

Off-Policy Average Reward Actor-Critic with Deterministic Policy Search

ICML 2023poster

The average reward criterion is relatively less studied as most existing works in the Reinforcement Learning literature consider the discounted reward criterion. There are few recent works that present on-policy average reward actor-critic algorithms, but average reward off-policy actor-critic is re…

Cited by 11SourcePDFScholar
2022

Dynamic Mirror Descent based Model Predictive Control for Accelerating Robot Learning

ICRA 2022poster

Recent works in Reinforcement Learning (RL) combine model-free (Mf)-RL algorithms with model-based (Mb)-RL approaches to get the best from both: asymptotic performance of Mf-RL and high sample-efficiency of Mb-RL. Inspired by these works, we propose a hierarchical framework that integrates online le…

Cited by 3SourceScholar
2022

Model-based Safe Deep Reinforcement Learning via a Constrained Proximal Policy Optimization Algorithm

NeurIPS 2022accept

During initial iterations of training in most Reinforcement Learning (RL) algorithms, agents perform a significant number of random exploratory steps. In the real world, this can limit the practicality of these algorithms as it can lead to potentially dangerous behavior. Hence safe exploration is a…

2020

Robust Quadrupedal Locomotion on Sloped Terrains: A Linear Policy Approach

CoRL 2020

In this paper, with a view toward fast deployment of locomotion gaits in low-cost hardware, we use a linear policy for realizing end-foot trajectories in the quadruped robot, Stoch 2. In particular, the parameters of the end-foot trajectories are shaped via a linear feedback policy that takes the to

2019

Realizing Learned Quadruped Locomotion Behaviors through Kinematic Motion Primitives

ICRA 2019poster

Humans and animals are believed to use a very minimal set of trajectories to perform a wide variety of tasks including walking. Our main objective in this paper is two fold 1) Obtain an effective tool to realize these basic motion patterns for quadrupedal walking, called the kinematic motion primiti…

Cited by 29SourceScholar