2026
PsumQuant: In-line Post-training Partial Sum Quantizer for Energy Efficient NPU Inference
ICML 2026poster
The rapid growth of deep neural networks (DNNs) has intensified the demand for efficient hardware acceleration under quantization. While prior research has successfully reduced weight and activation precision, partial sums generated during accumulation often retain high precision, resulting in signi…