← Search

Michael T Pearce

2 accepted papers

2025

Bilinear MLPs enable weight-based mechanistic interpretability

ICLR 2025spotlight

A mechanistic understanding of how MLPs do computation in deep neural net- works remains elusive. Current interpretability work can extract features from hidden activations over an input dataset but generally cannot explain how MLP weights construct features. One challenge is that element-wise nonli…

2025

Sparse Autoencoders Do Not Find Canonical Units of Analysis

ICLR 2025poster

A common goal of mechanistic interpretability is to decompose the activations of neural networks into features: interpretable properties of the input computed by the model. Sparse autoencoders (SAEs) are a popular method for finding these features in LLMs, and it has been postulated that they can be…

Cited by 1SourcePDFScholar