← Search

Vinayak Pathak

3 accepted papers

2026

Structural Inference: Interpreting Small Language Models with Susceptibilities

ICLR 2026poster

We develop a linear response framework for interpretability that treats a neural network as a Bayesian statistical mechanical system. A small perturbation of the data distribution, for example shifting the Pile toward GitHub or legal text, induces a first-order change in the posterior expectation of…

Cited by 0SourceScholar
2020

Black-box Certification and Learning under Adversarial Perturbations

ICML 2020poster

We formally study the problem of classification under adversarial perturbations from a learner’s perspective as well as a third-party who aims at certifying the robustness of a given black-box classifier. We analyze a PAC-type framework of semi-supervised learning and identify possibility and imposs…

Cited by 27SourcePDFScholar