← Search

Xinyi Gu

2 accepted papers

2026

Composition-Grounded Instruction Synthesis for Visual Reasoning

ICLR 2026poster

Pretrained multi-modal large language models (MLLMs) demonstrate strong performance on diverse multimodal tasks, but remain limited in reasoning capabilities for domains where annotations are difficult to collect. In this work, we focus on artificial image domains such as charts, rendered documents,…

Cited by 0SourcecodeScholar
2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

ICML 2026poster

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, overlooking realistic settings where numerical evidence in cha…

Cited by 0SourceScholar