2025
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA
EMNLP 2025
Large language models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks. However, as the model size and the input sequence’s length increase, the linearly increasing key-value (KV) cache significantly degrades inference throughput. Therefore, grouped-q