2024
Emergent Visual-Semantic Hierarchies in Image-Text Representations
ECCV 2024oral
"While recent vision-and-language models (VLMs) like CLIP are a powerful tool for analyzing text and images in a shared semantic space, they do not explicitly model the hierarchical nature of the set of texts which may describe an image. Conversely, existing multimodal hierarchical representation le…