2026
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
IJCAI 2026
Improving the memory efficiency, throughput, and serving cost of large language models (LLMs) is critical for edge deployment, interactive applications, and sustainable inference. Pruning is a promising approach, but existing methods have limitations: width pruning disrupts the standard transformer