2025
Enhancing Automated Interpretability with Output-Centric Feature Descriptions
ACL 2025long
Automated interpretability pipelines generate natural language descriptions for the concepts represented by features in large language models (LLMs), such as “plants” or “the first word in a sentence”. These descriptions are derived using inputs that activate the feature, which may be a dimension or…