← Search

Hyochan Chong

2 accepted papers

2026

NanoQuant: Efficient Sub-1-bit Quantization of Large Language Models

ICML 2026poster

Weight-only quantization has become a standard approach for efficiently serving large language models (LLMs). However, existing methods fail to efficiently compress models to binary (1-bit), as they either require large amounts of data and compute or incur additional storage. In this work, we propos…

Cited by 0SourceScholar
2026

RaBiT: Residual Aware Binarization Training for Accurate and Efficient LLMs

ICML 2026poster

Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization promises hardware-friendly, matmul-free inference by stacking binary ($\pm$1) layers, but is plagued by pathological feat…

Cited by 0SourceScholar