← Search

Philip J. Ball

6 accepted papers

2023

Efficient Online Reinforcement Learning with Offline Data

ICML 2023poster

Sample efficiency and exploration remain major challenges in online reinforcement learning (RL). A powerful approach that can be applied to address these issues is the inclusion of offline data, such as prior trajectories from a human expert or a sub-optimal exploration policy. Previous methods have…

2022

Learning General World Models in a Handful of Reward-Free Deployments

NeurIPS 2022accept

Building generally capable agents is a grand challenge for deep reinforcement learning (RL). To approach this challenge practically, we outline two key desiderata: 1) to facilitate generalization, exploration should be task agnostic; 2) to facilitate scalability, exploration policies should collect…

2022

Stabilizing Off-Policy Deep Reinforcement Learning from Pixels

ICML 2022spotlight

Off-policy reinforcement learning (RL) from pixel observations is notoriously unstable. As a result, many successful algorithms must combine different domain-specific practices and auxiliary losses to learn meaningful behaviors in complex environments. In this work, we provide novel analysis demonst…

2021

Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment

ICML 2021spotlight

Reinforcement learning from large-scale offline datasets provides us with the ability to learn policies without potentially unsafe or impractical exploration. Significant progress has been made in the past few years in dealing with the challenge of correcting for differing behavior between the data…

Cited by 56SourcePDFScholar
2019

The Sensitivity of Counterfactual Fairness to Unmeasured Confounding

UAI 2019poster

Causal approaches to fairness have seen substantial recent interest, both from the machine learning community and from wider parties interested in ethical prediction algorithms. In no small part, this has been due to the fact that causal models allow one to simultaneously leverage data and expert kn…