2025
FlexiGPT: Pruning and Extending Large Language Models with Low-Rank Weight Sharing
NAACL 2025long
The rapid proliferation of large language models (LLMs) in natural language processing (NLP) has created a critical need for techniques that enable efficient deployment on memory-constrained devices without compromising performance. We present a method to prune LLMs that selectively prunes model blo…