2025
MoBA: Mixture of Block Attention for Long-Context LLMs
NeurIPS 2025spotlight
Scaling the effective context length is essential for advancing large language models (LLMs) toward artificial general intelligence (AGI). However, the quadratic increase in computational complexity inherent in traditional attention mechanisms presents a prohibitive overhead. Existing approaches eit…