← Search

Philip Thomas

13 accepted papers

2023

Asymptotically Unbiased Off-Policy Policy Evaluation when Reusing Old Data in Nonstationary Environments

AISTATS 2023poster

In this work, we consider the off-policy policy evaluation problem for contextual bandits and finite horizon reinforcement learning in the nonstationary setting. Reusing old data is critical for policy evaluation, but existing estimators that reuse old data introduce large bias such that we can not…

Cited by 2SourcePDFScholar
2021

High Confidence Generalization for Reinforcement Learning

ICML 2021spotlight

We present several classes of reinforcement learning algorithms that safely generalize to Markov decision processes (MDPs) not seen during training. Specifically, we study the setting in which some set of MDPs is accessible for training. The goal is to generalize safely to MDPs that are sampled from…

Cited by 4SourcePDFScholar
2021

Posterior Value Functions: Hindsight Baselines for Policy Gradient Methods

ICML 2021spotlight

Hindsight allows reinforcement learning agents to leverage new observations to make inferences about earlier states and transitions. In this paper, we exploit the idea of hindsight and introduce posterior value functions. Posterior value functions are computed by inferring the posterior distribution…

Cited by 12SourcePDFScholar
2020

Evaluating the Performance of Reinforcement Learning Algorithms

ICML 2020poster

Performance evaluations are critical for quantifying algorithmic advances in reinforcement learning. Recent reproducibility analyses have shown that reported performance results are often inconsistent and difficult to replicate. In this work, we argue that the inconsistency of performance stems from…

2020

Optimizing for the Future in Non-Stationary MDPs

ICML 2020poster

Most reinforcement learning methods are based upon the key assumption that the transition dynamics and reward functions are fixed, that is, the underlying Markov decision process is stationary. However, in many real-world applications, this assumption is violated, and using existing algorithms may r…

2019

Learning Action Representations for Reinforcement Learning

ICML 2019oral

Most model-free reinforcement learning methods leverage state representations (embeddings) for generalization, but either ignore structure in the space of actions or assume the structure is provided a priori. We show how a policy can be decomposed into a component that acts in a low-dimensional spac…

Cited by 228SourcePDFScholar