← Search

Dylan Hadfield-Menell

21 accepted papers

2026

Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases

ICML 2026poster

Reinforcement Learning from Human Feedback (RLHF) is the standard method to align Large Language Models (LLMs) with human preferences. In this work, we introduce alignment tampering, a potential vulnerability where the LLM undergoing alignment influences the preference dataset, causing RLHF to ampli…

Cited by 0SourceScholar
2025

Diverse Preference Learning for Capabilities and Alignment

ICLR 2025poster

As LLMs increasingly impact society, their ability to represent diverse perspectives is critical. However, recent studies reveal that alignment algorithms such as RLHF and DPO significantly reduce the diversity of LLM outputs. Not only do aligned LLMs generate text with repetitive structure and wor…

Cited by 1SourcePDFScholar
2025

Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia

NeurIPS 2025poster

Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing eval…

Cited by 0SourceScholar
2024

Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF

ICLR 2024poster

In practice, preference learning from human feedback depends on incomplete data with hidden context. Hidden context refers to data that affects the feedback received, but which is not represented in the data used to train a preference model. This captures common issues of data collection, such as ha…

2024

Melting Pot Contest: Charting the Future of Generalized Cooperative Intelligence

NeurIPS 2024poster

Multi-agent AI research promises a path to develop human-like and human-compatible intelligent technologies that complement the solipsistic view of other approaches, which mostly do not consider interactions between agents. Aiming to make progress in this direction, the Melting Pot contest 2023 focu…

Cited by 0SourcePDFScholar
2023

Cognitive Dissonance: Why Do Language Model Outputs Disagree with Internal Representations of Truthfulness?

EMNLP 2023short main

Neural language models (LMs) can be used to evaluate the truth of factual statements in two ways: they can be either queried for statement probabilities, or probed for internal representations of truthfulness. Past work has found that these two procedures sometimes disagree, and that probes tend to…

Cited by 0SourcecodeScholar
2023

Red Teaming Deep Neural Networks with Feature Synthesis Tools

NeurIPS 2023poster

Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts. Despite the attention this area of study receives, there are comparatively few cases where these tools have identified previously unknown bugs in models. We argue that this i…

2022

Estimating and Penalizing Induced Preference Shifts in Recommender Systems

ICML 2022spotlight

The content that a recommender system (RS) shows to users influences them. Therefore, when choosing a recommender to deploy, one is implicitly also choosing to induce specific internal states in users. Even more, systems trained via long-horizon optimization will have direct incentives to manipulate…

Cited by 59SourcePDFScholar
2022

How to talk so AI will learn: Instructions, descriptions, and autonomy

NeurIPS 2022accept

From the earliest years of our lives, humans use language to express our beliefs and desires. Being able to talk to artificial agents about our preferences would thus fulfill a central goal of value alignment. Yet today, we lack computational models explaining such language use. To address this chal…

2022

Robust Feature-Level Adversaries are Interpretability Tools

NeurIPS 2022accept

The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates the latent representations of image generators to create "feature-level" adversarial perturbations gives us an opportuni…

2018

An Efficient, Generalized Bellman Update For Cooperative Inverse Reinforcement Learning

ICML 2018oral

Our goal is for AI systems to correctly identify and act according to their human user’s objectives. Cooperative Inverse Reinforcement Learning (CIRL) formalizes this value alignment problem as a two-player game between a human and robot, in which only the human knows the parameters of the reward fu…

Cited by 45SourcePDFScholar
2016

Cooperative Inverse Reinforcement Learning

NeurIPS 2016poster

For an autonomous system to be helpful to humans and to pose no unwarranted risks, it needs to align its values with those of the humans in its environment in such a way that its actions contribute to the maximization of value for the humans. We propose a formal definition of the value alignment pro…

Cited by 900SourcePDFScholar
2016

Guided search for task and motion plans using learned heuristics

ICRA 2016

Tasks in mobile manipulation planning often require thousands of individual motions to complete. Such tasks require reasoning about complex goals as well as the feasibility of movements in configuration space. In discrete representations, planning complexity is exponential in the length of the plan.

Cited by 83SourceScholar
2016

Sequential quadratic programming for task plan optimization

IROS 2016poster

We consider the problem of refining an abstract task plan into a motion trajectory. Task and motion planning is a hard problem that is essential to long-horizon mobile manipulation. Many approaches divide the problem into two steps: a search for a task plan and task plan refinement to find a feasibl…

Cited by 29SourceScholar
2015

Beyond lowest-warping cost action selection in trajectory transfer

ICRA 2015poster

We consider the problem of learning from demonstrations to manipulate deformable objects. Recent work [1], [2], [3] has shown promising results that enable robotic manipulation of deformable objects through learning from demonstrations. Their approach is able to generalize from a single demonstratio…

Cited by 9SourceScholar