← Search

Kyungphil Park

2 accepted papers

2024

Towards Next-Level Post-Training Quantization of Hyper-Scale Transformers

NeurIPS 2024poster

With the increasing complexity of generative AI models, post-training quantization (PTQ) has emerged as a promising solution for deploying hyper-scale models on edge devices such as mobile and TVs. Existing PTQ schemes, however, consume considerable time and resources, which could be a bottleneck in…

2023

A Frustratingly Easy Post-Training Quantization Scheme for LLMs

EMNLP 2023long main

Efficient inference has become crucial for hyper-scale AI models, including large language models, as their parameter count continues to increase for enhanced performance. This necessity holds true regardless of the computing environment, whether it be mobile devices or cloud servers. Quantization e…

Cited by 0SourceScholar