Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions are often vague, inconsistent, and require manual relabeling. In response, we introduce *semantic regexes*, structured language desc…