ACL 2025long0 citations

Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models

Jie Liu, Wenxuan Wang, Su Yihang, Jingyuan Huang, Yudi Zhang, Cheng-Yi Li, Wenting Chen, Xiaohan Xing

Abstract

The significant breakthroughs of Medical Multi-Modal Large Language Models (Med-MLLMs) renovate modern healthcare with robust information synthesis and medical decision support. However, these models are often evaluated on benchmarks that are unsuitable for the Med-MLLMs due to the intricate nature of the real-world diagnostic frameworks, which encompass diverse medical specialties and involve complex clinical decisions. Thus, a clinically representative benchmark is highly desirable for credible Med-MLLMs evaluation. To this end, we introduce Asclepius, a novel Med-MLLM benchmark that comprehensively assesses Med-MLLMs in terms of: distinct medical specialties (cardiovascular, gastroenterology, etc.) and different diagnostic capacities (perception, disease analysis, etc.). Grounded in 3 proposed core principles, Asclepius ensures a comprehensive evaluation by encompassing 15 medical specialties, stratifying into 3 main categories and 8 sub-categories of clinical tasks, and exempting overlap with the existing VQA dataset. We further provide an in-depth analysis of 6 Med-MLLMs and compare them with 3 human specialists, providing insights into their competencies and limitations in various medical contexts. Our work not only advances the understanding of Med-MLLMs’ capabilities but also sets a precedent for future evaluations and the safe deployment of these models in clinical environments.

BibTeX
@inproceedings{liu-etal-2025-asclepius,
    title = "Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models",
    author = "Liu, Jie  and
      Wang, Wenxuan  and
      Yihang, Su  and
      Huang, Jingyuan  and
      Zhang, Yudi  and
      Li, Cheng-Yi  and
      Chen, Wenting  and
      Xing, Xiaohan  and
      Chang, Kao-Jung  and
      Shen, Linlin  and
      Lyu, Michael R.",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1178/",
    doi = "10.18653/v1/2025.acl-long.1178",
    pages = "24181--24201",
    ISBN = "979-8-89176-251-0"
}
Asclepius: A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models · ACL 2025