2025
ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
ICML 2025spotlight
Post-training quantization (PTQ) of large language models (LLMs) holds the promise in reducing the prohibitive computational cost at inference time. Quantization of all weight, activation and key-value (KV) cache tensors to 4-bit without significantly degrading generalizability is challenging, due t…