Adversary-Aware DPO: Enhancing Safety Alignment in Vision Language Models via Adversarial Training
Safety alignment is critical in pre-trained large language models (LLMs) to generate responses aligned with human values and refuse harmful queries. Unlike LLM, the current safety alignment of VLMs is often achieved with post-hoc safety fine-tuning. However, these methods are less effective to white