2024
Multi-modal Preference Alignment Remedies Degradation of Visual Instruction Tuning on Language Models
ACL 2024long
Multi-modal large language models (MLLMs) are expected to support multi-turn queries of interchanging image and text modalities in production. However, the current MLLMs trained with visual-question-answering (VQA) datasets could suffer from degradation, as VQA datasets lack the diversity and comple…