2026
D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness
AAAI 2026technical
Large language models (LLMs) face significant deployment challenges due to their massive computational demands. While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and t