NAACL 2025findings13 citations

MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Zayd Muhammad Kawakibi Zuhri, Muhammad Farid Adilazuarda, Ayu Purwarianti, Alham Fikri Aji

Abstract

Auto-regressive inference of transformers benefit greatly from Key-Value (KV) caching, but can lead to major memory bottlenecks as model size, batch size, and sequence length grow at scale. We introduce Multi-Layer Key-Value (MLKV) sharing, a novel approach extending KV sharing across transformer layers to reduce memory usage beyond what was possible with Multi-Query Attention (MQA) and Grouped-Query Attention (GQA). Evaluations on various NLP benchmarks and inference metrics using uptrained Pythia-160M variants demonstrate that MLKV significantly reduces memory usage with minimal performance loss, reducing KV cache size down to a factor of 6x compared to MQA. These results highlight MLKV’s potential for efficient deployment of transformer models at scale.

BibTeX
@inproceedings{zuhri-etal-2025-mlkv,
    title = "{MLKV}: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding",
    author = "Zuhri, Zayd Muhammad Kawakibi  and
      Adilazuarda, Muhammad Farid  and
      Purwarianti, Ayu  and
      Aji, Alham Fikri",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Findings of the Association for Computational Linguistics: NAACL 2025",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-naacl.305/",
    pages = "5516--5525",
    ISBN = "979-8-89176-195-7"
}
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding · NAACL 2025