Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign Relearning
Machine unlearning is a promising approach to mitigate undesirable memorization of training data in ML models. However, in this work we show that existing approaches for unlearning in LLMs are surprisingly susceptible to a simple set of benign relearning attacks. With access to only a small and pote…