2026
SAEmnesia: Erasing Concepts in Diffusion Models with Supervised Sparse Autoencoders
ICML 2026poster
Concept unlearning in diffusion models is hampered by feature splitting, where concepts are distributed across many latent features, making their removal challenging and computationally expensive. We introduce SAEmnesia, a supervised sparse autoencoder framework that overcomes this by enforcing one-…