Benchmarking the Scientific Mind: Toward Evaluation of Complex-Reasoning Biomedical VQA
Despite progress of Multimodal Large Language Models (MLLMs) in biomedical visual question answering (VQA), existing benchmarks provide limited assessment of their scientific reasoning capabilities. Most datasets adopt single-image question construction and outcome-oriented evaluation, where correct…