2026
DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
ICML 2026poster
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, overlooking realistic settings where numerical evidence in cha…