← Search

Anish Mudide

1 accepted papers

2025

Efficient Dictionary Learning with Switch Sparse Autoencoders

ICLR 2025poster

Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features represented in frontier models, it will be necessary to scale them up to very high width, posing a computational challen…