← Search

Ali Edalati

2 accepted papers

2025

OAC: Output-adaptive Calibration for Accurate Post-training Quantization

AAAI 2025technical

Deployment of Large Language Models (LLMs) has major computational costs, due to their rapidly expanding size. Compression of LLMs reduces the memory footprint, latency, and energy required for their inference. Post-training Quantization (PTQ) techniques have been developed to compress LLMs while a…

Cited by 0SourcePDFScholar
2022

Kronecker Decomposition for GPT Compression

ACL 2022short

GPT is an auto-regressive Transformer-based pre-trained language model which has attracted a lot of attention in the natural language processing (NLP) domain. The success of GPT is mostly attributed to its pre-training on huge amount of data and its large number of parameters. Despite the superior p…

Cited by 43SourcePDFScholar