Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman, Matthew Jagielski, Katja Filippova, Ken Liu, Alexandra Chouldechova
Abstract
"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact.
BibTeX
@inproceedings{
cooper2025machine,
title={Machine Unlearning Doesn't Do What You Think: Lessons for Generative {AI} Policy and Research},
author={A. Feder Cooper and Christopher A. Choquette-Choo and Miranda Bogen and Kevin Klyman and Matthew Jagielski and Katja Filippova and Ken Liu and Alexandra Chouldechova and Jamie Hayes and Yangsibo Huang and Eleni Triantafillou and Peter Kairouz and Nicole Elyse Mitchell and Niloofar Mireshghallah and Abigail Z. Jacobs and James Grimmelmann and Vitaly Shmatikov and Christopher De Sa and Ilia Shumailov and Andreas Terzis and Solon Barocas and Jennifer Wortman Vaughan and danah boyd and Yejin Choi and Sanmi Koyejo and Fernando Delgado and Percy Liang and Daniel E. Ho and Pamela Samuelson and Miles Brundage and David Bau and Seth Neel and Hanna Wallach and Amy B. Cyphert and Mark Lemley and Nicolas Papernot and Katherine Lee},
booktitle={The Thirty-Ninth Annual Conference on Neural Information Processing Systems Position Paper Track},
year={2025},
url={https://openreview.net/forum?id=mfd6GRW4Az}
}