← Search

Chi-Chun Liu

2 accepted papers

2026

Is Finer Better? The Limits of Microscaling Formats in Large Language Models

ICLR 2026poster

Microscaling data formats leverage per-block tensor quantization to enable aggressive model compression with limited loss in accuracy. Unlocking their potential for efficient training and inference necessitates hardware-friendly implementations that handle matrix multiplications in a native format a…

Cited by 0SourcecodeScholar
2022

Deep Compression of Pre-trained Transformer Models

NeurIPS 2022accept

Pre-trained transformer models have achieved remarkable success in natural language processing (NLP) and have recently become competitive alternatives to Convolution Neural Networks (CNN) and Recurrent Neural Networks (RNN) in vision and speech tasks, respectively. Due to excellent computational eff…

Cited by 22SourcePDFScholar