2025
Seeing Beyond: Enhancing Visual Question Answering with Multi-Modal Retrieval
COLING 2025industry
Multi-modal Large language models (MLLMs) have made significant strides in complex content understanding and reasoning. However, they still suffer from model hallucination and lack of specific knowledge when facing challenging questions. To address these limitations, retrieval augmented generation (…