2024
A2Q+: Improving Accumulator-Aware Weight Quantization
ICML 2024poster
Quantization techniques commonly reduce the inference costs of neural networks by restricting the precision of weights and activations. Recent studies show that also reducing the precision of the accumulator can further improve hardware efficiency at the risk of numerical overflow, which introduces…