← Search

Banseok Lee

3 accepted papers

2026

RaBiT: Residual Aware Binarization Training for Accurate and Efficient LLMs

ICML 2026poster

Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance. Residual binarization promises hardware-friendly, matmul-free inference by stacking binary ($\pm$1) layers, but is plagued by pathological feat…

Cited by 0SourceScholar
2025

LittleBit: Ultra Low-Bit Quantization via Latent Factorization

NeurIPS 2025poster

Deploying large language models (LLMs) often faces challenges from substantial memory and computational costs. Quantization offers a solution, yet performance degradation in the sub-1-bit regime remains particularly difficult. This paper introduces LittleBit, a novel method for extreme LLM compressi…

Cited by 0SourceScholar