2026
PolySAE: Modeling Feature Interactions in Sparse Autoencoders via Polynomial Decoding
ICML 2026poster
Sparse autoencoders (SAEs) have emerged as a promising method for interpreting neural network representations by decomposing activations into sparse combinations of dictionary atoms. However, SAEs assume that features combine additively through linear reconstruction, an assumption that cannot captur…