2026
Resting Neurons, Active Insights: Robustify Activation Sparsity for Large Language Models
ICML 2026poster
Activation sparsity offers a compelling route to accelerate large language model (LLM) inference by selectively suppressing hidden activations, yet existing approaches exhibit severe accuracy degradation at high sparsity. We show that this failure stems from representational instability: *activation…