SkillCLIP: Skill Aware Modality Fusion Visual Question Answering (Student Abstract)
When humans are posed with a difficult problem, they often approach it by identifying key skills, honing them, and finally effectively combining them. We propose a novel method and apply it for the VizWiz VQA task to predict the visual skills needed to answer a question, and leverage expert modules…