← Search

Sidharth Baskaran

2 accepted papers

2025

HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks

ICLR 2025poster

Mechanistic interpretability has made great strides in identifying neural network features (e.g., directions in hidden activation space) that mediate concepts (e.g., *the birth year of a Nobel laureate*) and enable predictable manipulation. Distributed alignment search (DAS) leverages supervision fr…

Cited by 0SourcePDFScholar
2024

Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing

EMNLP 2024main

Recent work using Rank-One Model Editing (ROME), a popular model editing method, has shown that there are certain facts that the algorithm is unable to edit without breaking the model. Such edits have previously been called disabling edits. These disabling edits cause immediate model collapse and li…