2024
Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering
EMNLP 2024main
While large pre-trained visual-language models have shown promising results on traditional visual question answering benchmarks, it is still challenging for them to answer complex VQA problems which requires diverse world knowledge. Motivated by the research of retrieval-augmented generation in the…