← Search

Ahmet Caner Yüzügüler

2 accepted papers

2026

TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix

ICLR 2026poster

Multi-Head Latent Attention (MLA) is a recent attention mechanism adopted in state-of-the-art LLMs such as DeepSeek-v3 and Kimi K2. Thanks to its novel formulation, MLA allows two functionally equivalent but computationally distinct kernel implementations: naive and absorb. While the naive kernels (…

Cited by 0SourceScholar
2022

U-Boost NAS: Utilization-Boosted Differentiable Neural Architecture Search

ECCV 2022poster

"Optimizing resource utilization in target platforms is key to achieving high performance during DNN inference. While optimizations have been proposed for inference latency, memory footprint, and energy consumption, prior hardware-aware neural architecture search (NAS) methods have omitted resource…