2025
Analyzing (In)Abilities of SAEs via Formal Languages
NAACL 2025long
Autoencoders have been used for finding interpretable and disentangled features underlying neural network representations in both image and text domains. While the efficacy and pitfalls of such methods are well-studied in vision, there is a lack of corresponding results, both qualitative and quantit…