2025
IG-Pruning: Input-Guided Block Pruning for Large Language Models
EMNLP 2025
With the growing computational demands of large language models (LLMs), efficient inference has become increasingly critical for practical deployment. Depth pruning has emerged as a promising approach for reducing the computational costs of large language models by removing transformer layers. Howev