← Search

Sokratis Trifinopoulos

1 accepted papers

2024

From Neurons to Neutrons: A Case Study in Interpretability

ICML 2024poster

Mechanistic Interpretability (MI) proposes a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and…