← Search

Jacob Coxon

1 accepted papers

2026

Weight-sparse transformers have interpretable circuits

ICML 2026poster

Finding human-understandable circuits in language models is a central goal of the field of mechanistic interpretability. We train models to have more understandable circuits by constraining most of their weights to be zeros, so that each neuron only has a few connections. To recover fine-grained cir…

Cited by 0SourceScholar