← Search

Xiaoyao Liang

2 accepted papers

2023

$\rm A^2Q$: Aggregation-Aware Quantization for Graph Neural Networks

ICLR 2023poster

As graph data size increases, the vast latency and memory consumption during inference pose a significant challenge to the real-world deployment of Graph Neural Networks (GNNs). While quantization is a powerful approach to reducing GNNs complexity, most previous works on GNNs quantization fail to ex…

2021

Improving Neural Network Efficiency via Post-Training Quantization With Adaptive Floating-Point

ICCV 2021poster

Model quantization has emerged as a mandatory technique for efficient inference with advanced Deep Neural Networks (DNN). It converts the model parameters in full precision (32-bit floating point) to the hardware friendly data representation with shorter bit-width, to not only reduce the model size…

Cited by 58PDFcodeScholar