Group-SAE: Efficient Training of Sparse Autoencoders for Large Language Models via Layer Groups
Sparse AutoEncoders (SAEs) have recently been employed as a promising unsupervised approach for understanding the representations of layers of Large Language Models (LLMs). However, with the growth in model size and complexity, training SAEs is computationally intensive, as typically one SAE is trai