2022
F8Net: Fixed-Point 8-bit Only Multiplication for Network Quantization
ICLR 2022oral
Neural network quantization is a promising compression technique to reduce memory footprint and save energy consumption, potentially leading to real-time inference. However, there is a performance gap between quantized and full-precision models. To reduce it, existing quantization approaches require…