ICML 2026oral0 citations

Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections

Lukasz Borchmann, Jordy Van Landeghem, Michał Turski, Shreyansh Padarha, Ryan Kearns, Adam Mahdi, Niels Rogge, Clémentine Fourrier

Abstract

Multimodal agents offer a compelling path to automating complex document-intensive workflows, yet a critical question remains: do these architectures demonstrate genuine strategic reasoning, or simply conduct stochastic trial-and-error search? To address this, we introduce Agentic Document VQA, a benchmark of 2,250 human-authored questions grounded in 800 heterogeneous PDF documents. Guided by *Classical Test Theory*, we design it to maximize discriminative power and reliably differentiate between varying levels of agent capability. To rigorously assess agentic behaviour, we introduce a novel evaluation protocol for measuring the accuracy-effort trade-off. Using this framework, we find that humans show strong metacognitive calibration, adapting or abandoning failed strategies, whereas frontier agents often persist in unproductive loops with diminishing returns. We release the dataset, evaluation harness, and leaderboard to help facilitate the transition from brute-force retrieval to calibrated, efficient reasoning.

AgentsMultimodalRetrievalBenchmarkRobotics
BibTeX
@inproceedings{
borchmann2026strategic,
title={Strategic Navigation or Stochastic Search? How Agents and Humans Reason Over Document Collections},
author={{\L}ukasz Borchmann and Jordy Van Landeghem and Micha{\l} Turski and Shreyansh Padarha and Ryan Othniel Kearns and Adam Mahdi and Niels Rogge and Cl{\'e}mentine Fourrier and Siwei Han and Huaxiu Yao and Artemis Llabr{\'e}s and Yiming Xu and Dimosthenis Karatzas and Hao Zhang and Anupam Datta},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=ds3ZOevkwx}
}