← Search

Yilun Luo

1 accepted papers

2026

MicroMix: Efficient Mixed-Precision Quantization with Microscaling Formats for Large Language Models

ICLR 2026poster

Quantization significantly accelerates inference in large language models (LLMs) by replacing original high-precision matrices with low-precision counterparts. Recent advances in weight-activation quantization have primarily focused on mapping both weights and activations to the INT4 format. Althoug…

Cited by 0SourcecodeScholar