2026
SERQ: Saliency-Aware Low-Rank Error Reconstruction for LLM Quantization
ICLR 2026poster
Post-training quantization (PTQ) has emerged as a prevailing technique for deploying large language models (LLMs) efficiently in terms of both memory and computation, across edge devices and server platforms. Existing PTQ methods primarily aim to reduce precision in weights and activations by mitiga…