← Search

Arvind Satyanarayan

3 accepted papers

2026

Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language

ICLR 2026poster

Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions are often vague, inconsistent, and require manual relabeling. In response, we introduce *semantic regexes*, structured language desc…

Cited by 0SourcecodeScholar
2022

Teaching Humans When to Defer to a Classifier via Exemplars

AAAI 2022technical

Expert decision makers are starting to rely on data-driven automated agents to assist them with various tasks. For this collaboration to perform properly, the human decision maker must have a mental model of when and when not to rely on the agent. In this work, we aim to ensure that human decision m…