← Search

Zhengling Qi

14 accepted papers

2025

A Principled Path to Fitted Distributional Evaluation

NeurIPS 2025spotlight

In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy. This work focuses on extending the widely used fitted Q-evaluation---developed for expectation-based reinforce…

Cited by 0SourcecodeScholar
2025

Distributional Off-policy Evaluation with Bellman Residual Minimization

AISTATS 2025poster

We study distributional off-policy evaluation (OPE), of which the goal is to learn the distribution of the return for a target policy using offline data generated by a different policy. The theoretical foundation of many existing work relies on the supremum-extended statistical distances such as sup…

Cited by 0SourcecodeScholar
2024

A Fine-grained Analysis of Fitted Q-evaluation: Beyond Parametric Models

ICML 2024poster

In this paper, we delve into the statistical analysis of the fitted Q-evaluation (FQE) method, which focuses on estimating the value of a target policy using offline data generated by some behavior policy. We provide a comprehensive theoretical understanding of FQE estimators under both parametric a…

Cited by 0SourcePDFScholar
2024

Robust Offline Reinforcement Learning with Heavy-Tailed Rewards

AISTATS 2024poster

This paper endeavors to augment the robustness of offline reinforcement learning (RL) in scenarios laden with heavy-tailed rewards, a prevalent circumstance in real-world applications. We propose two algorithmic frameworks, ROAM and ROOM, for robust off-policy evaluation and offline policy optimizat…

2024

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

NeurIPS 2024poster

This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data literature, we propose a two-way unmeasured confounding assumption to model the system dynamics in causal reinforcement learn…

2023

Optimizing Pessimism in Dynamic Treatment Regimes: A Bayesian Learning Approach

AISTATS 2023poster

In this article, we propose a novel pessimism-based Bayesian learning method for optimal dynamic treatment regimes in the offline setting. When the coverage condition does not hold, which is common for offline data, the existing solutions would produce sub-optimal policies. The pessimism principle a…

2023

PASTA: Pessimistic Assortment Optimization

ICML 2023poster

We consider a fundamental class of assortment optimization problems in an offline data-driven setting. The firm does not know the underlying customer choice model but has access to an offline dataset consisting of the historically offered assortment set, customer choice, and revenue. The objective i…

Cited by 5SourcePDFScholar
2023

Pessimistic Model Selection for Offline Deep Reinforcement Learning

UAI 2023poster

Deep Reinforcement Learning (DRL) has demonstrated great potentials in solving sequential decision making problems in many applications. Despite its promising performance, practical gaps exist when deploying DRL in real-world scenarios. One main barrier is the over-fitting issue that leads to poor g…

Cited by 5SourcePDFScholar
2022

Off-Policy Evaluation for Episodic Partially Observable Markov Decision Processes under Non-Parametric Models

NeurIPS 2022accept

We study the problem of off-policy evaluation (OPE) for episodic Partially Observable Markov Decision Processes (POMDPs) with continuous states. Motivated by the recently proposed proximal causal inference framework, we develop a non-parametric identification result for estimating the policy value v…

Cited by 17SourcePDFScholar
2022

On Well-posedness and Minimax Optimal Rates of Nonparametric Q-function Estimation in Off-policy Evaluation

ICML 2022spotlight

We study the off-policy evaluation (OPE) problem in an infinite-horizon Markov decision process with continuous states and actions. We recast the $Q$-function estimation into a special form of the nonparametric instrumental variables (NPIV) estimation problem. We first show that under one mild condi…

Cited by 38SourcePDFScholar
2022

RISE: Robust Individualized Decision Learning with Sensitive Variables

NeurIPS 2022accept

This paper introduces RISE, a robust individualized decision learning framework with sensitive variables, where sensitive variables are collectible data and important to the intervention decision, but their inclusion in decision making is prohibited due to reasons such as delayed availability or fai…