2023
Self-Distilled Quantization: Achieving High Compression Rates in Transformer-Based Language Models
ACL 2023short
We investigate the effects of post-training quantization and quantization-aware training on the generalization of Transformer language models. We present a new method called self-distilled quantization (SDQ) that minimizes accumulative quantization errors and outperforms baselines. We apply SDQ to m…