2025
Wasserstein Distances, Neuronal Entanglement, and Sparsity
ICLR 2025spotlight
Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used to understand performance, particularly under weight sparsity, a leading post-training optimization technique. We sugges…