← Search

Patrick Shafto

13 accepted papers

2025

Convergence Theorems for Entropy-Regularized and Distributional Reinforcement Learning

NeurIPS 2025poster

In the pursuit of finding an optimal policy, reinforcement learning (RL) methods generally ignore the properties of learned policies apart from their expected return. Thus, even when successful, it is difficult to characterize which policies will be learned and what they will do. In this work, we pr…

Cited by 0SourceScholar
2024

Action Gaps and Advantages in Continuous-Time Distributional Reinforcement Learning

NeurIPS 2024poster

When decisions are made at high frequency, traditional reinforcement learning (RL) methods struggle to accurately estimate action values. In turn, their performance is inconsistent and often poor. Whether the performance of distributional RL (DRL) agents suffers similarly, however, is unknown. In th…

Cited by 1SourcePDFScholar
2021

Interactive Learning from Activity Description

ICML 2021spotlight

We present a novel interactive learning protocol that enables training request-fulfilling agents by verbally describing their activities. Unlike imitation learning (IL), our protocol allows the teaching agent to provide feedback in a language that is most appropriate for them. Compared with reward i…

2020

Interpretable Deep Gaussian Processes with Moments

AISTATS 2020poster

Deep Gaussian Processes (DGPs) combine the the expressiveness of Deep Neural Networks (DNNs) with quantified uncertainty of Gaussian Processes (GPs). Expressive power and intractable inference both result from the non-Gaussian distribution over composition functions. We propose interpretable DGP bas…

Cited by 21SourcePDFScholar
2018

Optimal Cooperative Inference

AISTATS 2018poster

Cooperative transmission of data fosters rapid accumulation of knowledge by efficiently combining experiences across learners. Although well studied in human learning and increasingly in machine learning, we lack formal frameworks through which we may reason about the benefits and limitations of coo…

Cited by 0SourcePDFScholar