2025
Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
EMNLP 2025
Rapid advances in Multimodal Large Language Models (MLLMs) have extended information retrieval beyond text, enabling access to complex real-world documents that combine both textual and visual content. However, most documents are private, either owned by individuals or confined within corporate silo