← Search

Puning Yang

3 accepted papers

2026

Distinguishable Deletion: Unifying Knowledge Erasure and Refusal for Large Language Model Unlearning

ICML 2026poster

Mitigating sensitive and harmful outputs is fundamental to ensuring safe deployment of LLMs. Existing approaches typically follow two paradigms: Knowledge Deletion (KD), which erases undesirable information during training, and Distinguishable Refusal (DR), which steers models away from using sensit…

Cited by 0SourceScholar
2025

Exploring Criteria of Loss Reweighting to Enhance LLM Unlearning

ICML 2025poster

Loss reweighting has shown significant benefits for machine unlearning with large language models (LLMs). However, their exact functionalities are left unclear and the optimal strategy remains an open question, thus impeding the understanding and improvement of existing methodologies. In this paper,…

2025

Towards Effective Evaluations and Comparisons for LLM Unlearning Methods

ICLR 2025poster

The imperative to eliminate undesirable data memorization underscores the significance of machine unlearning for large language models (LLMs). Recent research has introduced a series of promising unlearning methods, notably boosting the practical significance of the field. Nevertheless, adopting a p…

Cited by 0SourcePDFScholar