2026
LATMiX: Learnable Affine Transformations for Microscaling Quantization of LLMs
ICML 2026poster
Post-training quantization (PTQ) is a widely used approach for reducing the memory and compute costs of large language models (LLMs). Recent studies have shown that applying invertible transformations to activations can significantly improve quantization robustness by reducing activation outliers; h…