← Search

Daniel Murfet

6 accepted papers

2026

Structural Inference: Interpreting Small Language Models with Susceptibilities

ICLR 2026poster

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of…

Cited by 0SourceScholar
2026

Towards Spectroscopy: Susceptibility Clusters in Language Models

ICML 2026poster

Spectroscopy infers the internal structure of physical systems by measuring their response to perturbations. We apply this principle to neural networks: perturbing the data distribution by upweighting a token $y$ in context $x$, we measure the model's response via susceptibilities $\chi_{xy}$, which…

Cited by 0SourceScholar
2025

Differentiation and Specialization of Attention Heads via the Refined Local Learning Coefficient

ICLR 2025spotlight

We introduce refined variants of the Local Learning Coefficient (LLC), a measure of model complexity grounded in singular learning theory, to study the development of internal structure in transformer language models during training. By applying these refined LLCs (rLLCs) to individual components of…

Cited by 1SourcePDFScholar
2025

The Local Learning Coefficient: A Singularity-Aware Complexity Measure

AISTATS 2025poster

The Local Learning Coefficient (LLC) is introduced as a novel complexity measure for deep neural networks (DNNs). Recognizing the limitations of traditional complexity measures, the LLC leverages Singular Learning Theory (SLT), which has long recognized the significance of singularities in the loss…

Cited by 0SourcecodeScholar