2025
GuidedQuant: Large Language Model Quantization via Exploiting End Loss Guidance
ICML 2025poster
Post-training quantization is a key technique for reducing the memory and inference latency of large language models by quantizing weights and activations without requiring retraining. However, existing methods either (1) fail to account for the varying importance of hidden features to the end loss…