← Search

Alireza Ghaffari

2 accepted papers

2025

OAC: Output-adaptive Calibration for Accurate Post-training Quantization

AAAI 2025technical

Deployment of Large Language Models (LLMs) has major computational costs, due to their rapidly expanding size. Compression of LLMs reduces the memory footprint, latency, and energy required for their inference. Post-training Quantization (PTQ) techniques have been developed to compress LLMs while a…

Cited by 0SourcePDFScholar
2022

Is Integer Arithmetic Enough for Deep Learning Training?

NeurIPS 2022accept

The ever-increasing computational complexity of deep learning models makes their training and deployment difficult on various cloud and edge platforms. Replacing floating-point arithmetic with low-bit integer arithmetic is a promising approach to save energy, memory footprint, and latency of deep le…

Cited by 16SourcePDFScholar