← Search

Eric J Michaud

7 accepted papers

2025

Efficient Dictionary Learning with Switch Sparse Autoencoders

ICLR 2025poster

Sparse autoencoders (SAEs) are a recent technique for decomposing neural network activations into human-interpretable features. However, in order for SAEs to identify all features represented in frontier models, it will be necessary to scale them up to very high width, posing a computational challen…

2025

Not All Language Model Features Are One-Dimensionally Linear

ICLR 2025poster

Recent work has proposed that language models perform computation by manipulating one-dimensional representations of concepts ("features") in activation space. In contrast, we explore whether some language model representations may be inherently multi-dimensional. We begin by developing a rigorous d…

Cited by 0SourcePDFScholar
2025

On the creation of narrow AI: hierarchy and nonlocality of neural network skills

NeurIPS 2025poster

We study the problem of creating strong, yet narrow, AI systems. While recent AI progress has been driven by the training of large general-purpose foundation models, the creation of smaller models specialized for narrow domains could be valuable for both efficiency and safety. In this work, we explo…

Cited by 0SourcecodeScholar
2025

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

ICLR 2025oral

We introduce methods for discovering and applying **sparse feature circuits**. These are causally implicated subnetworks of human-interpretable features for explaining language model behaviors. Circuits identified in prior work consist of polysemantic and difficult-to-interpret units like attention…

2022

Towards Understanding Grokking: An Effective Theory of Representation Learning

NeurIPS 2022accept

We aim to understand grokking, a phenomenon where models generalize long after overfitting their training set. We present both a microscopic analysis anchored by an effective theory and a macroscopic analysis of phase diagrams describing learning performance across hyperparameters. We find that gene…