← Search

Chengzhuo Ni

6 accepted papers

2023

Representation Learning for Low-rank General-sum Markov Games

ICLR 2023poster

We study multi-agent general-sum Markov games with nonlinear function approximation. We focus on low-rank Markov games whose transition matrix admits a hidden low-rank structure on top of an unknown non-linear representation. The goal is to design an algorithm that (1) finds an $\varepsilon$-equilib…

Cited by 3SourcePDFScholar
2023

Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement

NeurIPS 2023poster

We explore the methodology and theory of reward-directed generation via conditional diffusion models. Directed generation aims to generate samples with desired properties as measured by a reward function, which has broad applications in generative AI, reinforcement learning, and computational biolog…

Cited by 36SourcePDFScholar
2022

Bandit Theory and Thompson Sampling-Guided Directed Evolution for Sequence Optimization

NeurIPS 2022accept

Directed Evolution (DE), a landmark wet-lab method originated in 1960s, enables discovery of novel protein designs via evolving a population of candidate sequences. Recent advances in biotechnology has made it possible to collect high-throughput data, allowing the use of machine learning to map out…

Cited by 6SourcePDFScholar
2022

Off-Policy Fitted Q-Evaluation with Differentiable Function Approximators: Z-Estimation and Inference Theory

ICML 2022spotlight

Off-Policy Evaluation (OPE) serves as one of the cornerstones in Reinforcement Learning (RL). Fitted Q Evaluation (FQE) with various function approximators, especially deep neural networks, has gained practical success. While statistical analysis has proved FQE to be minimax-optimal with tabular, li…

Cited by 23SourcePDFScholar
2022

Optimal Estimation of Policy Gradient via Double Fitted Iteration

ICML 2022spotlight

Policy gradient (PG) estimation becomes a challenge when we are not allowed to sample with the target policy but only have access to a dataset generated by some unknown behavior policy. Conventional methods for off-policy PG estimation often suffer from either significant bias or exponentially large…

Cited by 4SourcePDFScholar
2021

On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient Method

NeurIPS 2021spotlight

Policy gradient (PG) gives rise to a rich class of reinforcement learning (RL) methods. Recently, there has been an emerging trend to augment the existing PG methods such as REINFORCE by the \emph{variance reduction} techniques. However, all existing variance-reduced PG methods heavily rely on an u…

Cited by 85SourcePDFScholar