2024
WRP: Weight Recover Prune for Structured Sparsity
ACL 2024long
As the scale of Large Language Models (LLMs) increases, it is necessary to compress the models to reduce the substantial demand on computational resources. Network pruning significantly reduces the model size by converting the weight matrix from dense to sparse data format. Current methodologies adv…