← Search

Mike Williams

3 accepted papers

2024

From Neurons to Neutrons: A Case Study in Interpretability

ICML 2024poster

Mechanistic Interpretability (MI) proposes a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and…

2022

Towards Understanding Grokking: An Effective Theory of Representation Learning

NeurIPS 2022accept

We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set. We present both a microscopic analysis anchored by an effective theory and a macroscopic analysis of phase diagrams describing learning performance across hyperparameters. We find that gene…