2026
Rethinking Residual Errors in Compensation-based LLM Quantization
ICLR 2026poster
Methods based on weight compensation, which iteratively apply quantization and weight compensation to minimize the output error, have recently demonstrated remarkable success in quantizing Large Language Models (LLMs). The representative work, GPTQ, introduces several key techniques that make such…