AAAI 2026technical0 citations
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
Adeel Yousaf, Joseph Fioresi, James Beetham, Amrit Singh Bedi, Mubarak Shah
Abstract
Improving the safety of vision-language models like CLIP via fine-tuning often comes at a steep price, causing significant drops in their generalization performance. We find this trade-off stems from rigid alignment strategies that force unsafe concepts toward single, predefined safe targets, disrupting the model
BibTeX
@inproceedings{aaai2026_saferclipmitigat,
title = {SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge},
author = {Adeel Yousaf and Joseph Fioresi and James Beetham and Amrit Singh Bedi and Mubarak Shah},
booktitle = {AAAI 2026},
year = {2026}
}