← Search

Seung-taek Woo

2 accepted papers

2025

AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models

EMNLP 2025

To enable broader deployment of Large Language Models (LLMs), it is essential to identify the best-performing model under strict memory constraints. We present AMQ, Automated Mixed-Precision Weight-Only Quantization, a framework that assigns layer-wise quantization bit-widths to optimally balance mo