← Search

Binyan Zhang

2 accepted papers

2026

D2 Prune: Sparsifying Large Language Models via Dual Taylor Expansion and Attention Distribution Awareness

AAAI 2026technical

Large language models (LLMs) face significant deployment challenges due to their massive computational demands. While pruning offers a promising compression solution, existing methods suffer from two critical limitations: (1) They neglect activation distribution shifts between calibration data and t

Cited by 0SourcePDFScholar
2025

MPNAS: Multimodal Sentiment Analysis Pruning via Neural Architecture Search

ICASSP 2025accepted

With the rapid development of social media, sentiment analysis from multimodal posts has garnered significant attention in recent years. However, the substantial size of these models impedes their deployment on resource-constrained embedded devices. Although pruning has been extensively studied to r…

Cited by 0SourceScholar