AAAI 2026technical0 citations

MMIFEvol: Towards Evolutionary Multimodal Instruction Following

Haoyu Wang, Sihang Jiang, Xiangru Zhu, Yuyan Chen, Xiaojun Meng, Jiansheng Wei, Yitong Wang, Yanghua Xiao

Abstract

Multimodal Instruction Following serves as a fundamental capability of multimodal language models, involving accurate comprehension and execution of user-provided instructions. However, existing multimodal instruction-following datasets and benchmarks face the shortcomings outlined below: (a) Lack of Difficulty Stratification, they collect diverse instruction categories but neglect the stratification of difficulty levels across these categories, which leads to overlap, bias, and low interpretability. (b) Lack of Fine-Grained Metrics, they conflate the model

BibTeX
@inproceedings{aaai2026_mmifevoltowardse,
  title = {MMIFEvol: Towards Evolutionary Multimodal Instruction Following},
  author = {Haoyu Wang and Sihang Jiang and Xiangru Zhu and Yuyan Chen and Xiaojun Meng and Jiansheng Wei and Yitong Wang and Yanghua Xiao},
  booktitle = {AAAI 2026},
  year = {2026}
}
MMIFEvol: Towards Evolutionary Multimodal Instruction Following · AAAI 2026