← Search

Houjing Wei

1 accepted papers

2024

Find-the-Common: A Benchmark for Explaining Visual Patterns from Images

COLING 2024main

Recent advances in Instruction-fine-tuned Vision and Language Models (IVLMs), such as GPT-4V and InstructBLIP, have prompted some studies have started an in-depth analysis of the reasoning capabilities of IVLMs. However, Inductive Visual Reasoning, a vital skill for text-image understanding, remains…