2025
LongTableBench: Benchmarking Long-Context Table Reasoning across Real-World Formats and Domains
EMNLP 2025
We introduce LongTableBench , a benchmark for evaluating long-context reasoning over semi-structured tables across diverse formats, tasks, and domains. It comprises 5,950 QA instances spanning 7 table formats (e.g., Markdown, HTML, SQL), 18 domains, and input lengths up to 128K tokens, including mul