← Search

Yeo Jeong Park

1 accepted papers

2026

TurboBoA: Faster and Exact Attention-aware Quantization without Backpropagation

ICLR 2026poster

The rapid growth of large language models (LLMs) has heightened the importance of post-training quantization (PTQ) for reducing memory and computation costs. Among PTQ methods, GPTQ has gained significant attention for its efficiency, enabling billion-scale LLMs to be quantized within a few GPU hour…

Cited by 0SourcecodeScholar