← Search

Kenny Peng

6 accepted papers

2026

Position: Use Sparse Autoencoders to Discover Unknowns

ICML 2026poster

While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their usefulness. Here, we establish a conceptual distinction that reconciles competing narratives surrounding SAEs. We argue that even if SAEs may be less effective fo…

Cited by 0SourceScholar
2025

Sparse Autoencoders for Hypothesis Generation

ICML 2025poster

We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three steps: (1) train a sparse autoencoder on text embeddings to produce interpretable features describing the data distribu…

2024

Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers

NAACL 2024long

Large language models (LLMs) are dramatically influencing AI research, spurring discussions on what has changed so far and how to shape the field’s future. To clarify such questions, we analyze a new dataset of 16,979 LLM-related arXiv papers, focusing on recent trends in 2023 vs. 2018-2022. First,…