Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning
Mitigating sensitive and harmful outputs is fundamental to ensuring safe deployment of LLMs. Existing approaches typically follow two paradigms: Knowledge Deletion (KD), which erases undesirable information during training, and Distinguishable Refusal (DR), which steers models away from using sensit…