← Search

Navdeep Kumar

9 accepted papers

2026

Horizon Imagination: Efficient On-Policy Rollout in Diffusion World Models

ICLR 2026poster

We study diffusion-based world models for reinforcement learning, which offer high generative fidelity but face critical efficiency challenges in control. Current methods either require heavyweight models at inference or rely on highly sequential imagination, both of which impose prohibitive comput…

Cited by 0SourcecodeScholar
2025

Global Convergence of Policy Gradient in Average Reward MDPs

ICLR 2025poster

We present the first comprehensive finite-time global convergence analysis of policy gradient for infinite horizon average reward Markov decision processes (MDPs). Specifically, we focus on ergodic tabular MDPs with finite state and action spaces. Our analysis shows that the policy gradient iterates…

Cited by 0SourcePDFScholar
2025

Non-rectangular Robust MDPs with Normed Uncertainty Sets

NeurIPS 2025poster

Robust policy evaluation for non-rectangular uncertainty set is generally NP-hard, even in approximation. Consequently, existing approaches suffer from either exponential iteration complexity or significant accuracy gaps. Interestingly, we identify a powerful class of $L_p$-bounded uncertainty sets…

Cited by 0SourceScholar
2025

On the Convergence of Single-Timescale Actor-Critic

NeurIPS 2025poster

We analyze the global convergence of the single-timescale actor-critic (AC) algorithm for the infinite-horizon discounted Markov Decision Processes (MDPs) with finite state spaces. To this end, we introduce an elegant analytical framework for handling complex, coupled recursions inherent in the algo…

Cited by 0SourceScholar
2024

Bring Your Own (Non-Robust) Algorithm to Solve Robust MDPs by Estimating The Worst Kernel

ICML 2024poster

Robust Markov Decision Processes (RMDPs) provide a framework for sequential decision-making that is robust to perturbations on the transition kernel. However, current RMDP methods are often limited to small-scale problems, hindering their use in high-dimensional domains. To bridge this gap, we prese…

Cited by 1SourcePDFScholar
2024

Efficient Value Iteration for s-rectangular Robust Markov Decision Processes

ICML 2024poster

We focus on s-rectangular robust Markov decision processes (MDPs), which capture interconnected uncertainties across different actions within each state. This framework is more general compared to sa-rectangular robust MDPs, where uncertainties in each action are independent. However, the introduced…

Cited by 3SourcePDFScholar
2024

Solving Non-rectangular Reward-Robust MDPs via Frequency Regularization

AAAI 2024technical

In robust Markov decision processes (RMDPs), it is assumed that the reward and the transition dynamics lie in a given uncertainty set. By targeting maximal return under the most adversarial model from that set, RMDPs address performance sensitivity to misspecified environments. Yet, to preserve comp…

Cited by 2SourcePDFScholar
2023

Policy Gradient for Rectangular Robust Markov Decision Processes

NeurIPS 2023poster

Policy gradient methods have become a standard for training reinforcement learning agents in a scalable and efficient manner. However, they do not account for transition uncertainty, whereas learning robust policies can be computationally expensive. In this paper, we introduce robust policy gradient…

Cited by 36SourcePDFScholar
2022

The Geometry of Robust Value Functions

ICML 2022spotlight

The space of value functions is a fundamental concept in reinforcement learning. Characterizing its geometric properties may provide insights for optimization and representation. Existing works mainly focus on the value space for Markov Decision Processes (MDPs). In this paper, we study the geometry…

Cited by 8SourcePDFScholar