2025
Cost-Optimal Grouped-Query Attention for Long-Context Modeling
EMNLP 2025
Grouped-Query Attention (GQA) is a widely adopted strategy for reducing the computational cost of attention layers in large language models (LLMs). However, current GQA configurations are often suboptimal because they overlook how context length influences inference cost. Since inference cost grows