← Search

Yaqi Duan

10 accepted papers

2025

PILAF: Optimal Human Preference Sampling for Reward Modeling

ICML 2025poster

As large language models increasingly drive real-world applications, aligning them with human values becomes paramount. Reinforcement Learning from Human Feedback (RLHF) has emerged as a key technique, translating preference data into reward models when oracle human values remain inaccessible. In pr…

Cited by 1SourcePDFScholar
2024

Taming "data-hungry" reinforcement learning? Stability in continuous state-action spaces

NeurIPS 2024poster

We introduce a novel framework for analyzing reinforcement learning (RL) in continuous state-action spaces, and use it to prove fast rates of convergence in both off-line and on-line settings. Our analysis highlights two key stability properties, relating to how changes in value functions and/or pol…

Cited by 4SourcePDFScholar
2023

Invertible Residual Neural Networks with Conditional Injector and Interpolator for Point Cloud Upsampling

IJCAI 2023poster

Point clouds obtained by LiDAR and other sensors are usually sparse and irregular. Low-quality point clouds have serious influence on the final performance of downstream tasks. Recently, a point cloud upsampling network with normalizing flows has been proposed to address this problem. However, the n…

Cited by 2SourcePDFScholar
2022

Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with Pessimism

ICLR 2022poster

Offline reinforcement learning, which seeks to utilize offline/historical data to optimize sequential decision-making strategies, has gained surging prominence in recent studies. Due to the advantage that appropriate function approximators can help mitigate the sample complexity burden in modern rei…

Cited by 87SourcePDFScholar
2021

Bootstrapping Fitted Q-Evaluation for Off-Policy Inference

ICML 2021spotlight

Bootstrapping provides a flexible and effective approach for assessing the quality of batch reinforcement learning, yet its theoretical properties are poorly understood. In this paper, we study the use of bootstrapping in off-policy evaluation (OPE), and in particular, we focus on the fitted Q-evalu…

Cited by 52SourcePDFScholar
2021

Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient

ICML 2021spotlight

This paper provides a statistical analysis of high-dimensional batch reinforcement learning (RL) using sparse linear function approximation. When there is a large number of candidate features, our result sheds light on the fact that sparsity-aware methods can make batch RL more sample efficient. We…

Cited by 39SourcePDFScholar
2019

Learning low-dimensional state embeddings and metastable clusters from time series data

NeurIPS 2019poster

This paper studies how to find compact state embeddings from high-dimensional Markov state trajectories, where the transition kernel has a small intrinsic rank. In the spirit of diffusion map, we propose an efficient method for learning a low-dimensional state embedding and capturing the process's d…

Cited by 21SourcePDFScholar