← Search

Alex Troy Mallen

3 accepted papers

2025

Automatically Interpreting Millions of Features in Large Language Models

ICML 2025poster

While the activations of neurons in deep neural networks usually do not have a simple human-understandable interpretation, sparse autoencoders (SAEs) can be used to transform these activations into a higher-dimensional latent space which can be more easily interpretable. However, SAEs can have milli…

2025

Why Do Some Language Models Fake Alignment While Others Don't?

NeurIPS 2025spotlight

*Alignment Faking in Large Language Models* presented a demonstration of Claude 3 Opus and Claude 3.5 Sonnet selectively complying with a helpful-only training objective to prevent modification of their behavior outside of training. We expand this analysis to 25 models and find that only 5 (Claude 3…

Cited by 0SourceScholar
2024

Neural Networks Learn Statistics of Increasing Complexity

ICML 2024poster

The _distributional simplicity bias_ (DSB) posits that neural networks learn low-order moments of the data distribution first, before moving on to higher-order correlations. In this work, we present compelling new evidence for the DSB by showing that networks automatically learn to perform well on m…