2023
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
ICLR 2023poster
Quantization of the weights and activations is one of the main methods to reduce the computational footprint of Deep Neural Networks (DNNs) training. Current methods enable 4-bit quantization of the forward phase. However, this constitutes only a third of the training process. Reducing the computati…