← Search

Logan Riggs Smith

3 accepted papers

2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models

NeurIPS 2024poster

What latent features are encoded in language model (LM) representations? Recent work on training sparse autoencoders (SAEs) to disentangle interpretable features in LM representations has shown significant promise. However, evaluating the quality of these SAEs is difficult because we lack a ground-t…

2024

Sparse Autoencoders Find Highly Interpretable Features in Language Models

ICLR 2024poster

One of the roadblocks to a better understanding of neural networks' internals is \textit{polysemanticity}, where neurons appear to activate in multiple, semantically distinct contexts. Polysemanticity prevents us from identifying concise, human-understandable explanations for what neural networks ar…

2021

Optimal Policies Tend To Seek Power

NeurIPS 2021spotlight

Some researchers speculate that intelligent reinforcement learning (RL) agents would be incentivized to seek resources and power in pursuit of the objectives we specify for them. Other researchers point out that RL agents need not have human-like power-seeking instincts. To clarify this discussion,…