2026
Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models
ICML 2026poster
Sparse Autoencoders (SAEs) have become a cornerstone in mechanistic interpretability. However, current training methods inherit the Block Training paradigm from LLM pre-training. We identify this as a critical methodological oversight when applied to instruct models. Theoretically, utilizing GSNR an…