← Search

Philip S. Thomas

22 accepted papers

2026

Are Deep Speech Denoising Models Robust to Adversarial Noise?

ICLR 2026poster

Deep noise suppression (DNS) models enjoy widespread use throughout a variety of high-stakes speech applications. However, we show that four recent DNS models can each be reduced to outputting unintelligible gibberish through the addition of psychoacoustically hidden adversarial noise, even in low-…

Cited by 0SourceScholar
2025

Beyond Prediction: Managing the Repercussions of Machine Learning Applications

NeurIPS 2025poster

Machine learning models are often designed to maximize a primary goal, such as accuracy. However, as these models are increasingly used to inform decisions that affect people's lives or well-being, it is often unclear what the real-world repercussions of their deployment might be—making it crucial t…

Cited by 0SourceScholar
2025

Fair Continuous Resource Allocation with Equality of Impact

NeurIPS 2025poster

Recent works have studied fair resource allocation in social settings, where fairness is judged by the impact of allocation decisions rather than more traditional minimum or maximum thresholds on the allocations themselves. Our work significantly adds to this literature by developing continuous reso…

Cited by 0SourceScholar
2025

Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial Inference

NeurIPS 2025poster

Representation learning is increasingly applied to generate representations that generalize well across multiple downstream tasks. Ensuring fairness guarantees in representation learning is crucial to prevent unfairness toward specific demographic groups in downstream tasks. In this work, we formal…

Cited by 0SourceScholar
2024

Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation

NeurIPS 2024poster

Evaluating policies using off-policy data is crucial for applying reinforcement learning to real-world problems such as healthcare and autonomous driving. Previous methods for *off-policy evaluation* (OPE) generally suffer from high variance or irreducible bias, leading to unacceptably high predicti…

2024

From Past to Future: Rethinking Eligibility Traces

AAAI 2024technical

In this paper, we introduce a fresh perspective on the challenges of credit assignment and policy evaluation. First, we delve into the nuances of eligibility traces and explore instances where their updates may result in unexpected credit assignment to preceding states. From this investigation emerg…

Cited by 1SourcePDFScholar
2024

Position: Benchmarking is Limited in Reinforcement Learning Research

ICML 2024poster

Novel reinforcement learning algorithms, or improvements on existing ones, are commonly justified by evaluating their performance on benchmark environments and are compared to an ever-changing set of standard algorithms. However, despite numerous calls for improvements, experimental practices contin…

Cited by 7SourcePDFScholar
2023

Behavior Alignment via Reward Function Optimization

NeurIPS 2023spotlight

Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identification of reward structures that are not sparse and that avoid inadvertently inducing undesirable behaviors. Naively mo…

Cited by 15SourcePDFScholar
2022

Fairness Guarantees under Demographic Shift

ICLR 2022poster

Recent studies have demonstrated that using machine learning for social applications can lead to injustice in the form of racist, sexist, and otherwise unfair and discriminatory outcomes. To address this challenge, recent machine learning algorithms have been designed to limit the likelihood such un…

Cited by 64SourcePDFScholar
2022

Off-Policy Evaluation for Action-Dependent Non-stationary Environments

NeurIPS 2022accept

Methods for sequential decision-making are often built upon a foundational assumption that the underlying decision process is stationary. This limits the application of such methods because real-world problems are often subject to changes due to external factors (\textit{passive} non-stationarity),…

2021

High-Confidence Off-Policy (or Counterfactual) Variance Estimation

AAAI 2021technical

Many sequential decision-making systems leverage data collected using prior policies to propose a new policy. For critical applications, it is important that high-confidence guarantees on the new policy’s behavior are provided before deployment, to ensure that the policy will behave as desired. Prio…

Cited by 8SourcePDFScholar
2021

Multi-Objective SPIBB: Seldonian Offline Policy Improvement with Safety Constraints in Finite MDPs

NeurIPS 2021poster

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset collected under a known baseline policy, (ii) multiple reward signals are received from the environment inducing as many o…

Cited by 24SourcePDFScholar
2021

SOPE: Spectrum of Off-Policy Estimators

NeurIPS 2021poster

Many sequential decision making problems are high-stakes and require off-policy evaluation (OPE) of a new policy using historical data collected using some other policy. One of the most common OPE techniques that provides unbiased estimates is trajectory based importance sampling (IS). However, due…

2021

Structural Credit Assignment in Neural Networks using Reinforcement Learning

NeurIPS 2021poster

Structural credit assignment in neural networks is a long-standing problem, with a variety of alternatives to backpropagation proposed to allow for local training of nodes. One of the early strategies was to treat each node as an agent and use a reinforcement learning method called REINFORCE to upda…

Cited by 8SourcePDFScholar
2021

Universal Off-Policy Evaluation

NeurIPS 2021poster

When faced with sequential decision-making problems, it is often useful to be able to predict what would happen if decisions were made using a new policy. Those predictions must often be based on data collected under some previously used decision-making rule. Many previous methods enable such off-p…

2020

Security Analysis of Safe and Seldonian Reinforcement Learning Algorithms

NeurIPS 2020poster

We analyze the extent to which existing methods rely on accurate training data for a specific class of reinforcement learning (RL) algorithms, known as Safe and Seldonian RL. We introduce a new measure of security to quantify the susceptibility to perturbations in training data by creating an attack…

2020

Towards Safe Policy Improvement for Non-Stationary MDPs

NeurIPS 2020spotlight

Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that are safe for deployment, they assume that the underlying problem is stationary. However, many real-world problems of inte…

Cited by 32SourcePDFScholar
2019

A Meta-MDP Approach to Exploration for Lifelong Reinforcement Learning

NeurIPS 2019poster

In this paper we consider the problem of how a reinforcement learning agent that is tasked with solving a sequence of reinforcement learning problems (a sequence of Markov decision processes) can use knowledge acquired early in its lifetime to improve its ability to solve new problems. We argue that…

2019

Offline Contextual Bandits with High Probability Fairness Guarantees

NeurIPS 2019poster

We present RobinHood, an offline contextual bandit algorithm designed to satisfy a broad family of fairness constraints. Our algorithm accepts multiple fairness definitions and allows users to construct their own unique fairness definitions for the problem at hand. We provide a theoretical analysis of…

2017

Data-Efficient Policy Evaluation Through Behavior Policy Search

ICML 2017poster

We consider the task of evaluating a policy for a Markov decision process (MDP). The standard unbiased technique for evaluating a policy is to deploy the policy and observe its performance. We show that the data collected from deploying a different policy, commonly called the behavior policy, can be…

Cited by 55SourcePDFScholar
2017

Using Options and Covariance Testing for Long Horizon Off-Policy Policy Evaluation

NeurIPS 2017poster

Evaluating a policy by deploying it in the real world can be risky and costly. Off-policy policy evaluation (OPE) algorithms use historical data collected from running a previous policy to evaluate a new policy, which provides a means for evaluating a policy without requiring it to ever be deployed.…

Cited by 54SourcePDFScholar
2015

Policy Evaluation Using the Ω-Return

NeurIPS 2015poster

We propose the Ω-return as an alternative to the λ-return currently used by the TD(λ) family of algorithms. The benefit of the Ω-return is that it accounts for the correlation of different length returns. Because it is difficult to compute exactly, we suggest one way of approximating the Ω-return. W…

Cited by 21SourcePDFScholar