2026
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
ICLR 2026poster
Sparse Autoencoders (SAEs) are a powerful dictionary learning technique for decomposing neural network activations, translating the hidden state into human ideas with high semantic value despite no external intervention or guidance. However, this technique has rarely been applied outside of the text…