EMNLP 20250 citations
Probing Logical Reasoning of MLLMs in Scientific Diagrams
Abstract
We examine how multimodal large language models (MLLMs) perform logical inference grounded in visual information. We first construct a dataset of food web/chain images, along with questions that follow seven structured templates with progressively more complex reasoning involved. We show that complex reasoning about entities in the images remains challenging (even with elaborate prompts) and that visual information is underutilized.
BibTeX
@inproceedings{emnlp2025_probinglogicalre,
title = {Probing Logical Reasoning of MLLMs in Scientific Diagrams},
author = {Yufei Wang and Adriana Kovashka},
booktitle = {EMNLP 2025},
year = {2025}
}