2025
MuKA: Multimodal Knowledge Augmented Visual Information-Seeking
COLING 2025main
The visual information-seeking task aims to answer visual questions that require external knowledge, such as “On what date did this building officially open?”. Existing methods using retrieval-augmented generation framework primarily rely on textual knowledge bases to assist multimodal large languag…