AAAI 2026technical0 citations

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI

Tianbin Li, Yanzhou Su, Wei Li, Bin Fu, Zhe Chen, Ziyan Huang, Guoan Wang, Chenglong Ma

Abstract

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model

BibTeX
@inproceedings{aaai2026_gmaivlgmaivl55ma,
  title = {GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI},
  author = {Tianbin Li and Yanzhou Su and Wei Li and Bin Fu and Zhe Chen and Ziyan Huang and Guoan Wang and Chenglong Ma and Ying Chen and Ming Hu and Yanjun Li and Pengcheng Chen and Shixiang Tang and Xiaowei Hu and Zhongying Deng and Yuanfeng Ji and Jin Ye and Yu Qiao and Junjun He},
  booktitle = {AAAI 2026},
  year = {2026}
}
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI · AAAI 2026