2024
Safe-CLIP: Removing NSFW Concepts from Vision-and-Language Models
ECCV 2024poster
"Large-scale vision-and-language models, such as CLIP, are typically trained on web-scale data, which can introduce inappropriate content and lead to the development of unsafe and biased behavior. This, in turn, hampers their applicability in sensitive and trustworthy contexts and could raise signif…