← Search

Michael Sklar

1 accepted papers

2025

HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks

ICLR 2025poster

Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts (e.g., *the birth year of a Nobel laureate*) and enable predictable manipulation. Distributed alignment search (DAS) leverages supervision fr…

Cited by 0SourcePDFScholar