← Search

Jaehyun Jeon

2 accepted papers

2025

Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation

EMNLP 2025

Rapid advances in Multimodal Large Language Models (MLLMs) have extended information retrieval beyond text, enabling access to complex real-world documents that combine both textual and visual content. However, most documents are private, either owned by individuals or confined within corporate silo

Cited by 0SourcePDFScholar
2024

Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!

EMNLP 2024main

Humans possess multimodal literacy, allowing them to actively integrate information from various modalities to form reasoning. Faced with challenges like lexical ambiguity in text, we supplement this with other modalities, such as thumbnail images or textbook illustrations. Is it possible for machin…