NeurIPS 2025poster0 citations

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants

Hritik Bansal, Daniel Mingyi Israel, Siyan Zhao, Shufan Li, Tung Nguyen, Aditya Grover

Abstract

Recent advancements in mixed-modal generative have opened new avenues for developing unified biomedical assistants capable of analyzing biomedical images, answering complex questions about them, and generating multimodal patient reports. However, existing datasets face challenges such as small sizes, limited coverage of biomedical tasks and domains, and a reliance on narrow sources. To address these gaps, we present MedMax, a large-scale multimodal biomedical instruction-tuning dataset for mixed-modal foundation models. With 1.47 million instances, MedMax encompasses a diverse range of tasks, including interleaved image-text generation, biomedical image captioning and generation, visual chat, and report understanding. These tasks span knowledge across diverse biomedical domains, including radiology and histopathology, grounded in medical papers and YouTube videos. Subsequently, we fine-tune a mixed-modal foundation model on the MedMax dataset, achieving significant performance improvements: a 26% gain over the Chameleon model and an 18.3% improvement over GPT-4o across 12 downstream biomedical visual question-answering tasks. Finally, we introduce a unified evaluation suite for biomedical tasks to guide the development of mixed-modal biomedical AI assistants. We release the code, data, and model at https://mint-medmax.github.io/.

native multimodalbiomedical assistantinstruction tuning
BibTeX
@inproceedings{
bansal2025medmax,
title={MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants},
author={Hritik Bansal and Daniel Mingyi Israel and Siyan Zhao and Shufan Li and Tung Nguyen and Aditya Grover},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2025},
url={https://openreview.net/forum?id=mFEkBO25Ra}
}
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants · NeurIPS 2025