2024
Exploring the Benefit of Activation Sparsity in Pre-training
ICML 2024poster
Pre-trained Transformers inherently possess the characteristic of sparse activation, where only a small fraction of the neurons are activated for each token. While sparse activation has been explored through post-training methods, its potential in pre-training remains untapped. In this work, we firs…