Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM

Xiaoyu Wu (Rice University) · Yifei Pang (Carnegie Mellon University) · Terrance Liu (Carnegie Mellon University) · Steven Wu (Carnegie Mellon University)
data extraction attackexact unlearningextraction success rateslogits apimedical diagnosis datasetmodel guidancemuseopen-weight scenariosprivacy leakageprivacy risksthreat modelstofutoken filtering strategyunlearning methodswmdp

Large Language Models are typically trained on datasets collected from the web, which may inadvertently contain harmful or sensitive personal information. To address growing privacy concerns, unlearning methods have been proposed to remove the influence of specific data from trained models. Of these, exact unlearning