← Search

Jeremy Herbst

1 accepted papers

2026

The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level

ICML 2026poster

Mixture-of-Experts (MoE) architectures have become the dominant choice for scaling Large Language Models (LLMs), activating only a subset of parameters per token. While primarily adopted for computational efficiency, it remains an open question whether their sparsity makes them inherently easier to …

Cited by 0SourceScholar