2025
Monet: Mixture of Monosemantic Experts for Transformers
ICLR 2025poster
Understanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content generation. However, mechanistic interpretability is hindered by *polysemanticity*—where individual neurons respond to multip…