2022
Zero-Shot Dynamic Quantization for Transformer Inference
EMNLP 2022industry
We introduce a novel run-time method for significantly reducing the accuracy loss associated with quantizing BERT-like models to 8-bit integers. Existing methods for quantizing models either modify the training procedure, or they require an additional calibration step to adjust parameters that also…