← Search

Heejune Sheen

2 accepted papers

2026

Taming Polysemanticity in LLMs: Theory-Grounded Feature Recovery via Sparse Autoencoders

ICLR 2026poster

We study the challenge of achieving theoretically grounded feature recovery using Sparse Autoencoders (SAEs) for the interpretation of Large Language Models. Existing SAE training algorithms often lack rigorous mathematical guarantees and suffer from practical limitations such as hyperparameter sen…

Cited by 0SourcecodeScholar
2024

Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers

NeurIPS 2024poster

In-context learning (ICL) is a cornerstone of large language model (LLM) functionality, yet its theoretical foundations remain elusive due to the complexity of transformer architectures. In particular, most existing work only theoretically explains how the attention mechanism facilitates ICL under c…

Cited by 11SourcePDFScholar