Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM
data extraction attackexact unlearningextraction success rateslogits apimedical diagnosis datasetmodel guidancemuseopen-weight scenariosprivacy leakageprivacy risksthreat modelstofutoken filtering strategyunlearning methodswmdp
Large Language Models are typically trained on datasets collected from the web, which may inadvertently contain harmful or sensitive personal information. To address growing privacy concerns, unlearning methods have been proposed to remove the influence of specific data from trained models. Of these, exact unlearning