2025
Compress Large Language Models via Collaboration Between Learning and Matrix Approximation
NeurIPS 2025poster
Sparse and low-rank matrix composite approximation has emerged as a promising paradigm for compressing large language models (LLMs), offering a more flexible pruning structure than conventional methods based solely on sparse matrices. The significant variation in weight redundancy across layers, alo…