2026
DETAILCLIP: INJECTING IMAGE DETAILS INTO CLIP’S FEATURE SPACE
ICASSP 2026poster
Although CLIP-like Visual Language Models provide a functional joint feature space for image and text, due to the limitation of the CILP-like model's image input size (e.g., 224), subtle details are lost in the feature representation if we input high-resolution images (e.g., 2240). Our proposed fram…