← Search

Gonçalo Santos Paulo

1 accepted papers

2025

Automatically Interpreting Millions of Features in Large Language Models

ICML 2025poster

While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these activations into a higher-dimensional latent space which can be more easily interpretable. However, SAEs can have milli…