2025
On Localizing and Deleting Toxic Memories in Large Language Models
NAACL 2025findings
Warning: This paper contains offensive language.Ensuring that large language models (LLMs) do not generate harmful text is critical for their safe deployment. A common failure mode involves producing toxic responses to otherwise innocuous prompts. While various detoxification methods have been propo…