Prompting Large Language Models with Fine-Grained Visual Relations from Scene Graph for Visual Question Answering
Visual Question Answering (VQA) is a task that requires models to comprehend both questions and images. An increasing number of works are leveraging the strong reasoning capabilities of Large Language Models (LLMs) to address VQA. These methods typically utilize image captions as visual text descrip…