← Search

Andrew Perrault

18 accepted papers

2026

Optimizing Resource-Constrained Non-Pharmaceutical Interventions for Multi-Cluster Outbreak Control Using Hierarchical Reinforcement Learning

IJCAI 2026

Non-pharmaceutical interventions (NPIs), such as diagnostic testing and quarantine, are crucial for controlling infectious disease outbreaks but are often constrained by limited resources, particularly in the early outbreak stages. In real-world public health settings, resources must be allocated ac

Cited by 0Scholar
2025

Cultivating Archipelago of Forests: Evolving Robust Decision Trees Through Island Coevolution

AAAI 2025technical

Decision trees are widely used in machine learning due to their simplicity and interpretability, but they often lack robustness to adversarial attacks and data perturbations. The paper proposes a novel island-based coevolutionary algorithm (ICoEvoRDF) for constructing robust decision tree ensembles.…

2025

Symmetric Reinforcement Learning Loss for Robust Learning on Diverse Tasks and Model Scales

ICML 2025poster

Reinforcement learning (RL) training is inherently unstable due to factors such as moving targets and high gradient variance. Reinforcement Learning from Human Feedback (RLHF) and Reinforcement Learning from AI Feedback (RLAIF) introduce additional challenges. For instance, diverse preferences compl…

2025

The Distributional Reward Critic Framework for Reinforcement Learning Under Perturbed Rewards

AAAI 2025technical

The reward signal plays a central role in defining the desired behaviors of agents in reinforcement learning (RL). Rewards collected from realistic environments could be perturbed, corrupted, or noisy due to an adversary, sensor error, or because they come from subjective human feedback. Thus, it is…

2025

Using RLHF to align speech enhancement approaches to mean-opinion quality scores

ICASSP 2025accepted

Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and arti…

Cited by 8SourceScholar
2024

ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback

EMNLP 2024main

Large Multimodal Models (LMMs) excel at comprehending human instructions and demonstrate remarkable results across a broad spectrum of tasks. Reinforcement Learning from Human Feedback (RLHF) and AI Feedback (RLAIF) further refine LLMs by aligning them with specific preferences. These methods primar…

2024

Coevolutionary Algorithm for Building Robust Decision Trees under Minimax Regret

AAAI 2024technical

In recent years, there has been growing interest in developing robust machine learning (ML) models that can withstand adversarial attacks, including one of the most widely adopted, efficient, and interpretable ML algorithms—decision trees (DTs). This paper proposes a novel coevolutionary algorithm (…

2024

Leaving the Nest: Going beyond Local Loss Functions for Predict-Then-Optimize

AAAI 2024technical

Predict-then-Optimize is a framework for using machine learning to perform decision-making under uncertainty. The central research question it asks is, "How can we use the structure of a decision-making task to tailor ML models for that specific task?" To this end, recent work has proposed learning…

Cited by 14SourcePDFScholar
2022

Coordinating Followers to Reach Better Equilibria: End-to-End Gradient Descent for Stackelberg Games

AAAI 2022technical

A growing body of work in game theory extends the traditional Stackelberg game to settings with one leader and multiple followers who play a Nash equilibrium. Standard approaches for computing equilibria in these games reformulate the followers' best response as constraints in the leader's optimizat…

Cited by 31SourcePDFScholar
2022

Decision-Focused Learning without Decision-Making: Learning Locally Optimized Decision Losses

NeurIPS 2022accept

Decision-Focused Learning (DFL) is a paradigm for tailoring a predictive model to a downstream optimization task that uses its predictions in order to perform better \textit{on that specific task}. The main technical challenge associated with DFL is that it requires being able to differentiate throu…

Cited by 52SourcePDFScholar
2021

Learning MDPs from Features: Predict-Then-Optimize for Sequential Decision Making by Reinforcement Learning

NeurIPS 2021spotlight

In the predict-then-optimize framework, the objective is to train a predictive model, mapping from environment features to parameters of an optimization problem, which maximizes decision quality when the optimization is subsequently solved. Recent work on decision-focused learning shows that embeddi…

Cited by 38SourcePDFScholar
2021

Robust reinforcement learning under minimax regret for green security

UAI 2021poster

Green security domains feature defenders who plan patrols in the face of uncertainty about the adversarial behavior of poachers, illegal loggers, and illegal fishers. Importantly, the deterrence effect of patrols on adversaries’ future behavior makes patrol planning a sequential decision-making prob…

2020

Automatically Learning Compact Quality-aware Surrogates for Optimization Problems

NeurIPS 2020spotlight

Solving optimization problems with unknown parameters often requires learning a predictive model to predict the values of the unknown parameters and then solving the problem using these values. Recent work has shown that including the optimization problem as a layer in the model training pipeline re…

2020

Collapsing Bandits and Their Application to Public Health Intervention

NeurIPS 2020poster

We propose and study Collapsing Bandits, a new restless multi-armed bandit (RMAB) setting in which each arm follows a binary-state Markovian process with a special structure: when an arm is played, the state is fully observed, thus“collapsing” any uncertainty, but when an arm is passive, no observa…

2020

Robust Spatial-Temporal Incident Prediction

UAI 2020poster

Spatio-temporal incident prediction is a central issue in law enforcement, with applications in fighting crimes like poaching, human trafficking, illegal fishing, burglaries and smuggling. However, state of the art approaches fail to account for evasion in response to predictive models, a common fo…

Cited by 6SourcePDFScholar