ICLR 2026poster0 citations

Structural Inference: Interpreting Small Language Models with Susceptibilities

Garrett Baker, George Wang, Jesse Hoogland, Vinayak Pathak, Daniel Murfet

Abstract

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of an observable localized on a chosen component of the network. The resulting susceptibility can be estimated efficiently with local SGLD samples and factorizes into signed, per-token contributions that serve as attribution scores. We combine these susceptibilities into a response matrix whose low-rank structure separates functional modules such as multigram and induction heads in a 3M-parameter transformer.

InterpretabilityStatistical PhysicsSingular Learning Theory
BibTeX
@inproceedings{
baker2026structural,
title={Structural Inference: Interpreting Small Language Models with Susceptibilities},
author={Garrett Baker and George Wang and Jesse Hoogland and Vinayak Pathak and Daniel Murfet},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=J4GYMiE3JT}
}
Structural Inference: Interpreting Small Language Models with Susceptibilities · ICLR 2026