← Search

Arnob Ghosh

10 accepted papers

2025

Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity Guarantees

NeurIPS 2025poster

Constrained decision-making is essential for designing safe policies in real-world control systems, yet simulated environments often fail to capture real-world adversities. We consider the problem of learning a policy that will maximize the cumulative reward while satisfying a constraint, even when…

Cited by 0SourceScholar
2025

Provably Efficient RL for Linear MDPs under Instantaneous Safety Constraints in Non-Convex Feature Spaces

ICML 2025poster

In Reinforcement Learning (RL), tasks with instantaneous hard constraints present significant challenges, particularly when the decision space is non-convex or non-star-convex. This issue is especially relevant in domains like autonomous vehicles and robotics, where constraints such as collision avo…

Cited by 0SourcePDFScholar
2025

Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation

NeurIPS 2025spotlight

We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well u…

Cited by 0SourceScholar
2024

Adversarially Trained Weighted Actor-Critic for Safe Offline Reinforcement Learning

NeurIPS 2024poster

We propose WSAC (Weighted Safe Actor-Critic), a novel algorithm for Safe Offline Reinforcement Learning (RL) under functional approximation, which can robustly optimize policies to improve upon an arbitrary reference policy with limited data coverage. WSAC is designed as a two-player Stackelberg gam…

Cited by 0SourcePDFScholar
2024

Towards Achieving Sub-linear Regret and Hard Constraint Violation in Model-free RL

AISTATS 2024poster

We study the constrained Markov decision processes (CMDPs), in which an agent aims to maximize the expected cumulative reward subject to a constraint on the expected total value of a utility function. Existing approaches have primarily focused on \emph{soft} constraint violation, which allows compen…

Cited by 6SourcePDFScholar
2023

Achieving Sub-linear Regret in Infinite Horizon Average Reward Constrained MDP with Linear Function Approximation

ICLR 2023poster

We study the infinite horizon average reward constrained Markov Decision Process (CMDP). In contrast to existing works on model-based, finite state space, we consider the model-free linear CMDP setup. We first propose a computationally inefficient algorithm and show that $\tilde{\mathcal{O}}(\sqrt{…

Cited by 10SourcePDFScholar
2023

Provably Efficient Model-Free Algorithms for Non-stationary CMDPs

AISTATS 2023poster

We study model-free reinforcement learning (RL) algorithms in episodic non-stationary constrained Markov decision processes (CMDPs), in which an agent aims to maximize the expected cumulative reward subject to a cumulative constraint on the expected utility (cost). In the non-stationary environment,…

Cited by 21SourcePDFScholar
2022

Provably Efficient Model-Free Constrained RL with Linear Function Approximation

NeurIPS 2022accept

We study the constrained reinforcement learning problem, in which an agent aims to maximize the expected cumulative reward subject to a constraint on the expected total value of a utility function. In contrast to existing model-based approaches or model-free methods accompanied with a `simulator’,…

Cited by 36SourcePDFScholar