2024
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
NeurIPS 2024spotlight
In recent years, the input context sizes of large language models (LLMs) have increased dramatically. However, existing evaluation methods have not kept pace, failing to comprehensively assess the efficiency of models in handling long contexts. To bridge this gap, we introduce the BABILong benchmark…