2025
Computation and Memory-Efficient Model Compression with Gradient Reweighting
NeurIPS 2025poster
Pruning is a commonly employed technique for deep neural networks (DNNs) aiming at compressing the model size to reduce computational and memory costs during inference. In contrast to conventional neural networks, large language models (LLMs) pose a unique challenge regarding pruning efficiency due…