EMNLP 2024main5 citations

Dissecting Fine-Tuning Unlearning in Large Language Models

Yihuai Hong, Yuelin Zou, Lijie Hu, Ziqian Zeng, Di Wang, Haiqin Yang

Abstract

Fine-tuning-based unlearning methods prevail for erasing targeted harmful, sensitive, or copyrighted information within large language models while preserving overall capabilities. However, the true effectiveness of the methods is unclear. In this paper, we delve into the limitations of fine-tuning-based unlearning through activation patching and parameter restoration experiments. Our findings reveal that these methods alter the model’s knowledge retrieval process, rather than genuinely erasing the problematic knowledge embedded in the model parameters. Furthermore, behavioral tests demonstrate that the unlearning mechanisms inevitably impact the global behavior of the models, affecting unrelated knowledge or capabilities. Our work advocates the development of more resilient unlearning techniques for truly erasing knowledge.

BibTeX
@inproceedings{hong-etal-2024-dissecting,
    title = "Dissecting Fine-Tuning Unlearning in Large Language Models",
    author = "Hong, Yihuai  and
      Zou, Yuelin  and
      Hu, Lijie  and
      Zeng, Ziqian  and
      Wang, Di  and
      Yang, Haiqin",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.228/",
    doi = "10.18653/v1/2024.emnlp-main.228",
    pages = "3933--3941"
}
Dissecting Fine-Tuning Unlearning in Large Language Models · EMNLP 2024