← Search

Yaman Umuroglu

2 accepted papers

2026

MixQuant: Pushing the Limits of Block Rotations in Post-Training Quantization

ICML 2026poster

Recent post-training quantization (PTQ) methods have adopted block rotations to diffuse outliers prior to rounding. While this reduces the overhead of full-vector rotations, the effect of block structure on outlier suppression remains poorly understood. To fill this gap, we present the first systema…

Cited by 0SourceScholar
2024

A2Q+: Improving Accumulator-Aware Weight Quantization

ICML 2024poster

Quantization techniques commonly reduce the inference costs of neural networks by restricting the precision of weights and activations. Recent studies show that also reducing the precision of the accumulator can further improve hardware efficiency at the risk of numerical overflow, which introduces…

Cited by 5SourcePDFScholar