← Search

Cassidy Laidlaw

11 accepted papers

2025

AssistanceZero: Scalably Solving Assistance Games

ICML 2025poster

Assistance games are a promising alternative to reinforcement learning from human feedback (RLHF) for training AI assistants. Assistance games resolve key drawbacks of RLHF, such as incentives for deceptive behavior, by explicitly modeling the interaction between assistant and user as a two-player g…

2025

Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

ICLR 2025spotlight

Because it is difficult to precisely specify complex objectives, reinforcement learning policies are often optimized using proxy reward functions that only approximate the true goal. However, optimizing proxy rewards frequently leads to reward hacking: the optimized reward function ceases to be a go…

2025

Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision

ICLR 2025spotlight

Language model (LM) post-training relies on two stages of human supervision: task demonstrations for supervised finetuning (SFT), followed by preference comparisons for reinforcement learning from human feedback (RLHF). As LMs become more capable, the tasks they are given become harder to supervise.…

2024

Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

ICLR 2024poster

In practice, preference learning from human feedback depends on incomplete data with hidden context. Hidden context refers to data that affects the feedback received, but which is not represented in the data used to train a preference model. This captures common issues of data collection, such as ha…

2024

The Effective Horizon Explains Deep RL Performance in Stochastic Environments

ICLR 2024spotlight

Reinforcement learning (RL) theory has largely focused on proving minimax sample complexity bounds. These require strategic exploration algorithms that use relatively limited function classes for representing the policy or value function. Our goal is to explain why deep RL algorithms often perform w…

2023

Bridging RL Theory and Practice with the Effective Horizon

NeurIPS 2023oral

Deep reinforcement learning (RL) works impressively in some environments and fails catastrophically in others. Ideally, RL theory should be able to provide an understanding of why this is, i.e. bounds predictive of practical performance. Unfortunately, current theory does not quite have this ability…

2022

The Boltzmann Policy Distribution: Accounting for Systematic Suboptimality in Human Models

ICLR 2022poster

Models of human behavior for prediction and collaboration tend to fall into two categories: ones that learn from large amounts of data via imitation learning, and ones that assume human behavior to be noisily-optimal for some reward function. The former are very useful, but only when it is possible…

2021

Perceptual Adversarial Robustness: Defense Against Unseen Threat Models

ICLR 2021poster

A key challenge in adversarial robustness is the lack of a precise mathematical characterization of human perception, used in the definition of adversarial attacks that are imperceptible to human eyes. Most current attacks and defenses try to get around this issue by considering restrictive adversar…

2019

Capture, Learning, and Synthesis of 3D Speaking Styles

CVPR 2019poster

Audio-driven 3D facial animation has been widely explored, but achieving realistic, human-like performance is still unsolved. This is due to the lack of available 3D datasets, models, and standard evaluation metrics. To address this, we introduce a unique 4D face dataset with about 29 minutes of 4D…

Cited by 426PDFcodeScholar