AAAI 2026technical0 citations

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models

Jiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu, Yuzhuo Fu, Yangyang Kang

Abstract

Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models

BibTeX
@inproceedings{aaai2026_mmesciacomprehen,
  title = {MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models},
  author = {Jiacheng Ruan and Dan Jiang and Xian Gao and Ting Liu and Yuzhuo Fu and Yangyang Kang},
  booktitle = {AAAI 2026},
  year = {2026}
}
MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models · AAAI 2026