2025
UNDIAL: Self-Distillation with Adjusted Logits for Robust Unlearning in Large Language Models
NAACL 2025long
Mitigating the retention of sensitive or private information in large language models is essential for enhancing privacy and safety. Existing unlearning methods, like Gradient Ascent and Negative Preference Optimization, directly tune models to remove unwanted information. However, these methods oft…