← Search

Xiaoqing Sun

2 accepted papers

2026

Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit

ICML 2026poster

Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors. Current methods often rely on costly LLM-based techniques (e.g. annotating dataset differences) or dense embedding models (e.g. for clustering), which lack cont…

Cited by 0SourceScholar
2025

Dense SAE Latents Are Features, Not Bugs

NeurIPS 2025poster

Sparse autoencoders (SAEs) are designed to extract interpretable features from language models by enforcing a sparsity constraint. Ideally, training an SAE would yield latents that are both sparse and semantically meaningful. However, many SAE latents activate frequently (i.e., are *dense*), raising…

Cited by 0SourceScholar