2026
Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language
ICLR 2026poster
Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions are often vague, inconsistent, and require manual relabeling. In response, we introduce *semantic regexes*, structured language desc…