SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA over Text and Tables
Real-world Table–Text question answering (QA) tasks require models that can reason across long text and source tables, traversing multiple hops and executing complex operations such as aggregation. Yet existing benchmarks are small, manually curated—and therefore error-prone—and contain shallow ques…