ACL 2025long0 citations

OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

Xiangyu Zhao, Shengyuan Ding, Zicheng Zhang, Haian Huang, Maosongcao Maosongcao, Jiaqi Wang, Weiyun Wang, Xinyu Fang

Abstract

Recent advancements in open-source multi-modal large language models (MLLMs) have primarily focused on enhancing foundational capabilities, leaving a significant gap in human preference alignment. This paper introduces OmniAlign-V, a comprehensive dataset of 200K high-quality training samples featuring diverse images, complex questions, and varied response formats to improve MLLMs’ alignment with human preferences. We also present MM-AlignBench, a human-annotated benchmark specifically designed to evaluate MLLMs’ alignment with human values. Experimental results show that finetuning MLLMs with OmniAlign-V, using Supervised Fine-Tuning (SFT) or Direct Preference Optimization (DPO), significantly enhances human preference alignment while maintaining or enhancing performance on standard VQA benchmarks, preserving their fundamental capabilities.

BibTeX
@inproceedings{zhao-etal-2025-omnialign,
    title = "{O}mni{A}lign-{V}: Towards Enhanced Alignment of {MLLM}s with Human Preference",
    author = "Zhao, Xiangyu  and
      Ding, Shengyuan  and
      Zhang, Zicheng  and
      Huang, Haian  and
      Maosongcao, Maosongcao  and
      Wang, Jiaqi  and
      Wang, Weiyun  and
      Fang, Xinyu  and
      Wang, Wenhai  and
      Zhai, Guangtao  and
      Yang, Hua  and
      Duan, Haodong  and
      Chen, Kai",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.906/",
    doi = "10.18653/v1/2025.acl-long.906",
    pages = "18490--18515",
    ISBN = "979-8-89176-251-0"
}
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference · ACL 2025