2026
TyphoonMLA: A Mixed Naive-Absorb MLA Kernel For Shared Prefix
ICLR 2026poster
Multi-Head Latent Attention (MLA) is a recent attention mechanism adopted in state-of-the-art LLMs such as DeepSeek-v3 and Kimi K2. Thanks to its novel formulation, MLA allows two functionally equivalent but computationally distinct kernel implementations: naive and absorb. While the naive kernels (…