Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research

Yangsibo Huang (Google) · Jamie Hayes (Google DeepMind) · I Shumailov (University of Toronto) · Christopher Choquette-Choo (OpenAI) · Matthew Jagielski (Anthropic) · Niloofar Mireshghallah (UCSD) · Katherine Lee (OpenAI) · A. Feder Cooper (Stanford University) · Alex Chouldechova (Microsoft) · Solon Barocas (Microsoft Research; Cornell University) · Hanna Wallach (Microsoft) · Sanmi Koyejo (Stanford University / Virtue AI) · Miranda Bogen (Center for Democracy & Technology) · Kevin Klyman (Stanford University) · Katja Filippova (Research, Google) · Ken Liu (Stanford University) · Eleni Triantafillou (Google DeepMind) · Peter Kairouz (Google) · Nicole Mitchell (Google Research) · Abigail Jacobs (University of Michigan) · James Grimmelmann (Cornell University) · Vitaly Shmatikov (Cornell University) · Christopher De Sa (Cornell University) · Andreas Terzis (Google) · Jennifer Wortman Vaughan (Microsoft Research) · danah boyd (Data & Society Research Institute) · Yejin Choi (UW => Stanford / NVIDIA) · Fernando Delgado (Cornell University) · Percy Liang (Stanford University) · Daniel Ho (Stanford Law) · Pamela Samuelson (University of California, Berkeley) · Miles Brundage (OpenAI) · David Bau (Northeastern University) · Seth Neel (Google Research) · Amy Cyphert (West Virginia University) · Mark Lemley (Stanford University) · Nicolas Papernot (University of Toronto and Vector Institute)
behavioral circumscriptioncopyright issuesframework developmentgeneral-purpose solutiongenerative-aiinformation mitigationmachine unlearningml researchersmodel outputsmodel parameterspolicy implicationsprivacy concernssubstantive challengestargeted removaltargeted suppressiontechnical challenges

"Machine unlearning" is a popular proposed solution for mitigating the existence of content in an AI model that is problematic for legal or moral reasons, including privacy, copyright, safety, and more. For example, unlearning is often invoked as a solution for removing the effects of specific information from a generative-AI model's parameters, e.g., a particular individual's personal data or the inclusion of copyrighted content in the model's training data. Unlearning is also proposed as a way to prevent a model from generating targeted types of information in its outputs, e.g., generations that closely resemble a particular individual's data or reflect the concept of "Spiderman." Both of these goals--the targeted removal of information from a model and the targeted suppression of information from a model's outputs--present various technical and substantive challenges. We provide a framework for ML researchers and policymakers to think rigorously about these challenges, identifying several mismatches between the goals of unlearning and feasible implementations. These mismatches explain why unlearning is not a general-purpose solution for circumscribing generative-AI model behavior in service of broader positive impact.