AAAI 2024technical0 citations

SkillCLIP: Skill Aware Modality Fusion Visual Question Answering (Student Abstract)

Atharva Naik, Yash Parag Butala, Navaneethan Vaikunthan, Raghav Kapoor

Abstract

When humans are posed with a difficult problem, they often approach it by identifying key skills, honing them, and finally effectively combining them. We propose a novel method and apply it for the VizWiz VQA task to predict the visual skills needed to answer a question, and leverage expert modules to produce intermediary outputs and fuse them in a skill-aware manner. Unlike prior works in visual question-answering (VQA) that use intermediate outputs such as detected objects and Optical Character Recognition (OCR), our approach explicitly guides the model with a skill embedding on what to focus on. While our results show that using skill-aware fusion outperforms skill-unaware models for only a subset of questions, we believe our results provide interesting directions for future work. We also release our code, model, and illustrative demonstrations for future research purposes.

BibTeX
@article{Naik_Butala_Vaikunthan_Kapoor_2024, title={SkillCLIP: Skill Aware Modality Fusion Visual Question Answering (Student Abstract)}, volume={38}, url={https://ojs.aaai.org/index.php/AAAI/article/view/30486}, DOI={10.1609/aaai.v38i21.30486}, abstractNote={When humans are posed with a difficult problem, they often approach it by identifying key skills, honing them, and finally effectively combining them. We propose a novel method and apply it for the VizWiz VQA task to predict the visual skills needed to answer a question, and leverage expert modules to produce intermediary outputs and fuse them in a skill-aware manner. Unlike prior works in visual question-answering (VQA) that use intermediate outputs such as detected objects and Optical Character Recognition (OCR), our approach explicitly guides the model with a skill embedding on what to focus on. While our results show that using skill-aware fusion outperforms skill-unaware models for only a subset of questions, we believe our results provide interesting directions for future work. We also release our code, model, and illustrative demonstrations for future research purposes.}, number={21}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Naik, Atharva and Butala, Yash Parag and Vaikunthan, Navaneethan and Kapoor, Raghav}, year={2024}, month={Mar.}, pages={23592-23593} }