2025
SA-CLIP: Language Guided Image Spatial and Action Feature Learning
EMNLP 2025
We observed that Contrastive Language-Image Pretraining (CLIP) models struggle with real-world downstream tasks such as road traffic anomaly detection, due to their inability to effectively capture spatial and action relationships between objects within images. To address this, we compile and curate