2026
Feature Integration Spaces: Joint Training Reveals Dual Encoding in Neural Network Representations
AAAI 2026technical
Current sparse autoencoder (SAE) approaches to neural network interpretability assume that activations can be decomposed through linear superposition into sparse, interpret-able features. Despite high reconstruction fidelity, SAEs consistently fail to eliminate polysemanticity and exhibit pathologic