2025
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
EMNLP 2025
Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-quality and diverse training data. In this work, we aim to address the scarcity of reasoning-focused multimodal datasets.