← Search

Dannong Xu

3 accepted papers

2026

Do Vision and Text Cues Exhibit Evidential Coupling? UFO: A Benchmark for Compositional Multimodal Reasoning in Unified Models

ICML 2026poster

Unified Foundation Models (UFMs), which support interleaved multimodal generation and understanding, have been proposed as a promising paradigm for reasoning about dynamic world states, yet it remains unclear whether the visual content they generate functions as grounded evidence for subsequent reas…

Cited by 0SourceScholar
2025

Document Haystacks: Vision-Language Reasoning Over Piles of 1000+ Documents

CVPR 2025poster

Large multimodal models (LMMs) have achieved impressive progress in vision-language understanding, yet they face limitations in real-world applications requiring complex reasoning over a large number of images. Existing benchmarks for multi-image question-answering are limited in scope, each questio…

2025

WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation

ICCV 2025poster

Knowledge discovery and collection are intelligence-intensive tasks that traditionally require significant human effort to ensure high-quality outputs. Recent research has explored multi-agent frameworks for automating Wikipedia-style article generation by retrieving and synthesizing information fro…