← Search

Qiuhao Wang

4 accepted papers

2025

Provable Policy Gradient for Robust Average-Reward MDPs Beyond Rectangularity

ICML 2025poster

Robust Markov Decision Processes (MDPs) offer a promising framework for computing reliable policies under model uncertainty. While policy gradient methods have gained increasing popularity in robust discounted MDPs, their application to the average-reward criterion remains largely unexplored. This p…

Cited by 0SourcePDFScholar
2023

Global Convergence of Over-parameterized Deep Equilibrium Models

AISTATS 2023poster

A deep equilibrium model (DEQ) is implicitly defined through an equilibrium point of an infinite-depth weight-tied model with an input-injection. Instead of infinite computations, it solves an equilibrium point directly with root-finding and computes gradients with implicit differentiation. In this…

Cited by 16SourcePDFScholar
2023

Policy Gradient in Robust MDPs with Global Convergence Guarantee

ICML 2023poster

Robust Markov decision processes (RMDPs) provide a promising framework for computing reliable policies in the face of model errors. Many successful reinforcement learning algorithms build on variations of policy-gradient methods, but adapting these methods to RMDPs has been challenging. As a result,…