← Search

Ahmet Çelik

2 accepted papers

2026

SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights

ICML 2026poster

Post-training quantization has emerged as the most widely used strategy for deploying large language models at low precision. Still, current methods show perplexity degradation at bit-widths $\leq 4$, partly because representing outliers causes precision issues in parameters that share the same scal…

Cited by 5SourceScholar
2026

TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix

ICLR 2026poster

Multi-Head Latent Attention (MLA) is a recent attention mechanism adopted in state-of-the-art LLMs such as DeepSeek-v3 and Kimi K2. Thanks to its novel formulation, MLA allows two functionally equivalent but computationally distinct kernel implementations: naive and absorb. While the naive kernels (…

Cited by 0SourceScholar