← Search

Weilin Xu

1 accepted papers

2025

Soft Token Attacks Cannot Reliably Audit Unlearning in Large Language Models

EMNLP 2025

Large language models (LLMs) are trained using massive datasets.However, these datasets often contain undesirable content, e.g., harmful texts, personal information, and copyrighted material.To address this, machine unlearning aims to remove information from trained models.Recent work has shown that