2021
Neural gradients are near-lognormal: improved quantized and sparse training
ICLR 2021poster
While training can mostly be accelerated by reducing the time needed to propagate neural gradients (loss gradients with respect to the intermediate neural layer outputs) back throughout the model, most previous works focus on the quantization/pruning of weights and activations. These methods are oft…