2025
FLRC: Fine-grained Low-Rank Compressor for Efficient LLM Inference
EMNLP 2025
Although large language models (LLM) have achieved remarkable performance, their enormous parameter counts hinder deployment on resource-constrained hardware. Low-rank compression can reduce both memory usage and computational demand, but applying a uniform compression ratio across all layers often