D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness
Large language models (LLMs) face significant deployment challenges due to their massive computational demands. While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and t