At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization
Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diverges from training data distributions. Without explicit mechanisms for handling s…