2026
SEAP: Sparse Expert Activation Pruning Unlocks the Brainpower of Large Language Models
AAAI 2026technical
Pruning is a promising approach to reduce the high inference cost of large language models (LLMs), but it often comes at the expense of performance. Motivated by the "functional localization" theory in neuroscience, we hypothesize that LLMs contain task-specific expert activation paths, where specif