2024
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
EMNLP 2024main
The high power consumption and latency-sensitive deployments of large language models (LLMs) have motivated efficiency techniques like quantization and sparsity. Contextual sparsity, where the sparsity pattern is input-dependent, is crucial in LLMs because the permanent removal of attention heads or…