← Search

Alice Rigg

1 accepted papers

2025

Bilinear MLPs enable weight-based mechanistic interpretability

ICLR 2025spotlight

A mechanistic understanding of how MLPs do computation in deep neural net- works remains elusive. Current interpretability work can extract features from hidden activations over an input dataset but generally cannot explain how MLP weights construct features. One challenge is that element-wise nonli…