← Search

William Yeh

1 accepted papers

2025

Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders

EMNLP 2025

Large language models (LLMs) are now ubiquitous in user-facing applications, yet they still generate undesirable toxic outputs, including profanity, vulgarity, and derogatory remarks. Although numerous detoxification methods exist, most apply broad, surface-level fixes and can therefore easily be ci