2024
Exploring the Utility of Clip Priors for Visual Relationship Prediction
ICASSP 2024accepted
This work explores the challenges of leveraging large-scale vision language models, such as CLIP, for visual relationship prediction (VRP), a task vital in understanding the relations between objects in a scene based on both image features and text descriptors. Despite its potential, we find that CL…