← Search

Jakoba Petri-Koenig

2 accepted papers

2024

A2Q+: Improving Accumulator-Aware Weight Quantization

ICML 2024poster

Quantization techniques commonly reduce the inference costs of neural networks by restricting the precision of weights and activations. Recent studies show that also reducing the precision of the accumulator can further improve hardware efficiency at the risk of numerical overflow, which introduces…

Cited by 5SourcePDFScholar
2023

A2Q: Accumulator-Aware Quantization with Guaranteed Overflow Avoidance

ICCV 2023poster

We present accumulator-aware quantization (A2Q), a novel weight quantization method designed to train quantized neural networks (QNNs) to avoid overflow when using low-precision accumulators during inference. A2Q introduces a unique formulation inspired by weight normalization that constrains the L1…

Cited by 6PDFcodeScholar