ICML 2025poster0 citations

Prompt-based Depth Pruning of Large Language Models

Juyun Wee, Minjae Park, Jaeho Lee

Abstract

Depth pruning aims to reduce the inference cost of a large language model without any hardware-specific complications, by simply removing several less important transformer blocks. However, our empirical findings suggest that the importance of a transformer block may be highly task-dependent---a block that is crucial for a task can be removed without degrading the accuracy on another task. Based on this observation, we develop a dynamic depth pruning algorithm, coined PuDDing (**P**rompt-ro**u**ted **D**ynamic **D**epth Prun**ing**), which determines which blocks to omit from the model based on the input prompt. PuDDing operates by training a lightweight router to predict the best omission set among a set of options, where this option set has also been constructed in a data-driven manner. Empirical results on commonsense reasoning benchmarks demonstrate that PuDDing effectively accelerates the inference language models, and achieves better on-task performance than static depth pruning baselines.

Depth pruningModel Compression
BibTeX
@inproceedings{
wee2025promptbased,
title={Prompt-based Depth Pruning of Large Language Models},
author={Juyun Wee and Minjae Park and Jaeho Lee},
booktitle={Forty-second International Conference on Machine Learning},
year={2025},
url={https://openreview.net/forum?id=hRxHF1xPYB}
}
Prompt-based Depth Pruning of Large Language Models · ICML 2025