← Search

Matthew Zurek

6 accepted papers

2025

Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RL

NeurIPS 2025poster

We study offline reinforcement learning in average-reward MDPs, which presents increased challenges from the perspectives of distribution shift and non-uniform coverage, and has been relatively underexamined from a theoretical perspective. While previous work obtains performance guarantees under sin…

Cited by 0SourceScholar
2024

Learning to Stabilize Online Reinforcement Learning in Unbounded State Spaces

ICML 2024poster

In many reinforcement learning (RL) applications, we want policies that reach desired states and then keep the controlled system within an acceptable region around the desired states over an indefinite period of time. This latter objective is called *stability* and is especially important when the s…

2024

Span-Based Optimal Sample Complexity for Weakly Communicating and General Average Reward MDPs

NeurIPS 2024oral

We study the sample complexity of learning an $\varepsilon$-optimal policy in an average-reward Markov decision process (MDP) under a generative model. For weakly communicating MDPs, we establish the complexity bound $\widetilde{O}\left(SA\frac{\mathsf{H}}{\varepsilon^2} \right)$, where $\mathsf{H}$…

Cited by 3SourcePDFScholar
2023

The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types

AAAI 2023technical

When inferring reward functions from human behavior (be it demonstrations, comparisons, physical corrections, or e-stops), it has proven useful to model the human as making noisy-rational choices, with a "rationality coefficient" capturing how much noise or entropy we expect to see in the human beha…

Cited by 37SourcePDFScholar
2021

Situational Confidence Assistance for Lifelong Shared Autonomy

ICRA 2021poster

Shared autonomy enables robots to infer user intent and assist in accomplishing it. But when the user wants to do a new task that the robot does not know about, shared autonomy will hinder their performance by attempting to assist them with something that is not their intent. Our key idea is that th…

Cited by 34SourceScholar