L-CUBE: Isolating Long-Context Capacity from Knowledge with Controllable Mutual Information Scaling
Evaluating long-context language models on natural language conflates architectural capacity to capture dependencies with semantic knowledge and vocabulary statistics. When models fail at long contexts, we cannot determine whether failures stem from fundamental architectural limitations or insuffici…