← Search

Wenqian Zhao

4 accepted papers

2024

BiE: Bi-Exponent Block Floating-Point for Large Language Models Quantization

ICML 2024poster

Nowadays, Large Language Models (LLMs) mostly possess billions of parameters, bringing significant challenges to hardware platforms. Although quantization is an efficient approach to reduce computation and memory overhead for inference optimization, we stress the challenge that mainstream low-bit qu…

Cited by 5SourcePDFScholar
2024

Progressively Knowledge Distillation via Re-parameterizing Diffusion Reverse Process

AAAI 2024technical

Knowledge distillation aims at transferring knowledge from the teacher model to the student one by aligning their distributions. Feature-level distillation often uses L2 distance or its variants as the loss function, based on the assumption that outputs follow normal distributions. This poses a si…

Cited by 1SourcePDFScholar
2023

ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor Programs

EMNLP 2023long main

The training and inference efficiency of ever-larger deep neural networks highly rely on the performance of tensor operators on specific hardware platforms. Therefore, a compilation-based optimization flow with automatic tensor generation and parameter tuning is necessary for efficient model deploym…

Cited by 0SourceScholar
2021

Joint Semantic-geometric Learning for Polygonal Building Segmentation

AAAI 2021technical

Building extraction from aerial or satellite images has been an important research issue in remote sensing and computer vision domains for decades. Compared with pixel-wise semantic segmentation models that output raster building segmentation map, polygonal building segmentation approaches produce m…

Cited by 45SourcePDFScholar