2025
Retention Enhanced Cross-modal Attention for Multi-Hop VQA
ICASSP 2025accepted
Exploring multimodal information from external knowledge bases in Visual Question Answering (VQA) reasoning tasks presents a significant challenge. Current methods, such as attention-based and graph-based approaches, have limitations in effectively capturing contextual key information. It is essenti…