AAAI 2026technical0 citations

SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge

Adeel Yousaf, Joseph Fioresi, James Beetham, Amrit Singh Bedi, Mubarak Shah

Abstract

Improving the safety of vision-language models like CLIP via fine-tuning often comes at a steep price, causing significant drops in their generalization performance. We find this trade-off stems from rigid alignment strategies that force unsafe concepts toward single, predefined safe targets, disrupting the model

BibTeX
@inproceedings{aaai2026_saferclipmitigat,
  title = {SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge},
  author = {Adeel Yousaf and Joseph Fioresi and James Beetham and Amrit Singh Bedi and Mubarak Shah},
  booktitle = {AAAI 2026},
  year = {2026}
}
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge · AAAI 2026