2026
Learning multimodal dictionary decompositions with group-sparse autoencoders
ICLR 2026poster
The Linear Representation Hypothesis asserts that the embeddings learned by neural networks can be understood as linear combinations of features corresponding to high-level concepts. Based on this ansatz, sparse autoencoders (SAEs) have recently become a popular method for decomposing embeddings int…