DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
Jianyu Liu, Hangyu Guo, Ranjie Duan, Xingyuan Bu, Yancheng He, Shilong Li, Hui Huang, Jiaheng Liu
Abstract
Multimodal Large Language Models (MLLMs) pose unique safety challenges due to their integration of visual and textual data, thereby introducing new dimensions of potential attacks and complex risk combinations. In this paper, we begin with a detailed analysis aimed at disentangling risks through step-by-step reasoning within multimodal inputs. We find that systematic multimodal risk disentanglement substantially enhances the risk awareness of MLLMs. Via leveraging the strong discriminative abilities of multimodal risk disentanglement, we further introduce DREAM ( Disentangling Risks to Enhance Safety Alignment in MLLMs), a novel approach that enhances safety alignment in MLLMs through supervised fine-tuning and iterative Reinforcement Learning from AI Feedback (RLAIF). Experimental results show that DREAM significantly boosts safety during both inference and training phases without compromising performance on normal tasks (namely oversafety), achieving a 16.17% improvement in the SIUO safe&effective score compared to GPT-4V.
BibTeX
@inproceedings{liu-etal-2025-dream,
title = "{DREAM}: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models",
author = "Liu, Jianyu and
Guo, Hangyu and
Duan, Ranjie and
Bu, Xingyuan and
He, Yancheng and
Li, Shilong and
Huang, Hui and
Liu, Jiaheng and
Wang, Yucheng and
Jing, Chenchen and
Qu, Xingwei and
Zhang, Xiao and
Wang, Pei and
Wu, Yanan and
Gu, Jihao and
Li, Yangguang and
Zhu, Jianke",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-long.604/",
pages = "12097--12118",
ISBN = "979-8-89176-189-6"
}