2026
Pruning as a Cooperative Game: Surrogate-Assisted Layer Contribution Estimation for Large Language Models
ICLR 2026poster
While large language models (LLMs) demonstrate impressive performance across various tasks, their deployment in real-world scenarios is still constrained by high computational demands. Layer-wise pruning, a commonly employed strategy to mitigate inference costs, can partially address this challenge.…