2023
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
ICML 2023poster
Large language models (LLMs) show excellent performance but are compute- and memory-intensive. Quantization can reduce memory and accelerate inference. However, existing methods cannot maintain accuracy and hardware efficiency at the same time. We propose SmoothQuant, a training-free, accuracy-prese…