← Search

Ahn Young Jin

1 accepted papers

2025

Monet: Mixture of Monosemantic Experts for Transformers

ICLR 2025poster

Understanding the internal computations of large language models (LLMs) is crucial for aligning them with human values and preventing undesirable behaviors like toxic content generation. However, mechanistic interpretability is hindered by *polysemanticity*—where individual neurons respond to multip…