2026
Structural Inference: Interpreting Small Language Models with Susceptibilities
ICLR 2026poster
We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of…