AAAI 2026technical0 citations
MMIFEvol: Towards Evolutionary Multimodal Instruction Following
Haoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yitong Wang, Yanghua Xiao
Abstract
Multimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model
BibTeX
@inproceedings{aaai2026_mmifevoltowardse,
title = {MMIFEvol: Towards Evolutionary Multimodal Instruction Following},
author = {Haoyu Wang and Sihang Jiang and Xiangru Zhu and Yuyan Chen and Xiaojun Meng and Jiansheng Wei and Yitong Wang and Yanghua Xiao},
booktitle = {AAAI 2026},
year = {2026}
}