2024
Improved Image Captioning Via Knowledge Graph-Augmented Models
ICASSP 2024accepted
Multimodal foundation models, pre-trained on large-scale data, effectively capture vast amounts of factual and commonsense knowledge. However, these models store all their knowledge within their parameters, requiring increasingly larger models and training data to capture more knowledge. To address…