Budget-Aware LLM Quantization and Low-Rank Correction via Information-Guided Subspace Matrices
Compression techniques such as quantization and low-rank approximation enable large language models (LLMs) to run on current edge hardware with limited computing power, but the key challenge lies in balancing the allocation of precision and low rank within a fixed memory limit. We propose a training