← Search

Garrett Baker

2 accepted papers

2026

Structural Inference: Interpreting Small Language Models with Susceptibilities

ICLR 2026poster

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of…

Cited by 0SourceScholar
2026

Towards Spectroscopy: Susceptibility Clusters in Language Models

ICML 2026poster

Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token $y$ in context $x$, we measure the model's response via susceptibilities $\chi_{xy}$, which…

Cited by 0SourceScholar