2024
When Quantization Affects Confidence of Large Language Models?
NAACL 2024findings
Recent studies introduced effective compression techniques for Large Language Models (LLMs) via post-training quantization or low-bit weight representation. Although quantized weights offer storage efficiency and allow for faster inference, existing works have indicated that quantization might compr…