2026
Detoxifying Large Language Models via Localized Feature Editing with Sparse Autoencoders
IJCAI 2026
Large Language Models (LLMs) powerful generative capabilities also pose significant risks, underscoring the need for effective detoxification methods to ensure safer deployment. Due to the polysemantic nature of LLM neurons, recent neuron intervention methods inevitably entangle unrelated concepts,