A Benchmark for Deep Information Synthesis
Debjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas, Victor Prokhorov, Lena Sophia Bolliger, Aysim Toker, Roy Miles
Abstract
Large language model (LLM)-based agents are increasingly used to solve complex tasks involving tool use, such as web browsing, code execution, and data analysis. However, current evaluation benchmarks do not adequately assess their ability to solve real-world tasks that require synthesizing information from multiple sources and inferring insights beyond simple fact retrieval. To address this, we introduce DEEPSYNTH , a novel benchmark designed to evaluate agents on realistic, time-consuming problems that combine information gathering, synthesis, and structured reasoning to produce insights. DEEPSYNTH contains 120 tasks collected across 7 domains and data sources covering 42 countries. DEEPSYNTH is constructed using a multi-stage data collection pipeline that requires annotators to collect official data sources, create hypotheses, perform manual analysis and design tasks with verifiable answers. When evaluated on DEEPSYNTH, 9 state-of-the-art LLMs and deep research agents achieve a maximum F1 score of 8.97. Our analysis reveals that current agents struggle with hallucinations and reasoning over large information spaces, highlighting \ourdata as a crucial benchmark for guiding future research.
BibTeX
@inproceedings{
paul2026a,
title={A Benchmark for Deep Information Synthesis},
author={Debjit Paul and Daniel Murphy and Milan Gritta and Ronald Cardenas and Victor Prokhorov and Lena Sophia Bolliger and Aysim Toker and Roy Miles and Andreea-Maria Oncescu and Jasivan Alex Sivakumar and Philipp Borchert and Ismail Elezi and Meiru Zhang and Ka Yiu Lee and Guchun Zhang and Gerasimos Lampouras and Jun Wang},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=0Dhpt9aY3n}
}