2025
Prompt-based Depth Pruning of Large Language Models
ICML 2025poster
Depth pruning aims to reduce the inference cost of a large language model without any hardware-specific complications, by simply removing several less important transformer blocks. However, our empirical findings suggest that the importance of a transformer block may be highly task-dependent---a blo…