← Search

Roy Mayan

1 accepted papers

2025

Enhancing Automated Interpretability with Output-Centric Feature Descriptions

ACL 2025long

Automated interpretability pipelines generate natural language descriptions for the concepts represented by features in large language models (LLMs), such as “plants” or “the first word in a sentence”. These descriptions are derived using inputs that activate the feature, which may be a dimension or…