2022
Combining Compressions for Multiplicative Size Scaling on Natural Language Tasks
COLING 2022main
Quantization, knowledge distillation, and magnitude pruning are among the most popular methods for neural network compression in NLP. Independently, these methods reduce model size and can accelerate inference, but their relative benefit and combinatorial interactions have not been rigorously studie…