← Search

Yannick Assogba

2 accepted papers

2026

Semantic Regexes: Auto-Interpreting LLM Features with a Structured Language

ICLR 2026poster

Automated interpretability aims to translate large language model (LLM) features into human understandable descriptions. However, natural language feature descriptions are often vague, inconsistent, and require manual relabeling. In response, we introduce *semantic regexes*, structured language desc…

Cited by 0SourcecodeScholar
2022

Beyond Rewards: a Hierarchical Perspective on Offline Multiagent Behavioral Analysis

NeurIPS 2022accept

Each year, expert-level performance is attained in increasingly-complex multiagent domains, where notable examples include Go, Poker, and StarCraft II. This rapid progression is accompanied by a commensurate need to better understand how such agents attain this performance, to enable their safe depl…

Cited by 6SourcePDFScholar