2026
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
ICLR 2026poster
Post-training quantization (PTQ) compresses the weights and activations of large language models (LLMs) into low-precision representations to reduce memory footprint and accelerate inference. However, the presence of outliers in weights and activations often leads to large quantization errors and se…