← Search

James Wilken-Smith

1 accepted papers

2025

A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders

NeurIPS 2025oral

Sparse Autoencoders (SAEs) aim to decompose the activation space of large language models (LLMs) into human-interpretable latent directions or features. As we increase the number of features in the SAE, hierarchical features tend to split into finer features (“math” may split into “algebra”, “geomet…

Cited by 0SourcecodeScholar