← Search

Runzhe Wu

8 accepted papers

2025

Computationally Efficient RL under Linear Bellman Completeness for Deterministic Dynamics

ICLR 2025oral

We study computationally and statistically efficient Reinforcement Learning algorithms for the *linear Bellman Complete* setting. This setting uses linear function approximation to capture value functions and unifies existing models like linear Markov Decision Processes (MDP) and Linear Quadratic Re…

Cited by 5SourcePDFScholar
2025

Diffusing States and Matching Scores: A New Framework for Imitation Learning

ICLR 2025poster

Adversarial Imitation Learning is traditionally framed as a two-player zero-sum game between a learner and an adversarially chosen cost function, and can therefore be thought of as the sequential generalization of a Generative Adversarial Network (GAN). However, in recent years, diffusion models hav…

2023

Contextual Bandits and Imitation Learning with Preference-Based Active Queries

NeurIPS 2023poster

We consider the problem of contextual bandits and imitation learning, where the learner lacks direct knowledge of the executed action's reward. Instead, the learner can actively request the expert at each round to compare two actions and receive noisy preference feedback. The learner's objective is…

Cited by 21SourcePDFScholar
2023

Distributional Offline Policy Evaluation with Predictive Error Guarantees

ICML 2023poster

We study the problem of estimating the distribution of the return of a policy using an offline dataset that is not generated from the policy, i.e., distributional offline policy evaluation (OPE). We propose an algorithm called Fitted Likelihood Estimation (FLE), which conducts a sequence of Maximum…

2023

Selective Sampling and Imitation Learning via Online Regression

NeurIPS 2023poster

We consider the problem of Imitation Learning (IL) by actively querying noisy expert for feedback. While imitation learning has been empirically successful, much of prior work assumes access to noiseless expert feedback which is not practical in many applications. In fact, when one only has access t…

Cited by 9SourcePDFScholar
2023

The Benefits of Being Distributional: Small-Loss Bounds for Reinforcement Learning

NeurIPS 2023poster

While distributional reinforcement learning (DistRL) has been empirically effective, the question of when and why it is better than vanilla, non-distributional RL has remained unanswered. This paper explains the benefits of DistRL through the lens of small-loss bounds, which are instance-dependent b…

2021

Offline Constrained Multi-Objective Reinforcement Learning via Pessimistic Dual Value Iteration

NeurIPS 2021poster

In constrained multi-objective RL, the goal is to learn a policy that achieves the best performance specified by a multi-objective preference function under a constraint. We focus on the offline setting where the RL agent aims to learn the optimal policy from a given dataset. This scenario is common…

Cited by 20SourcePDFScholar