2022
Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective
AAAI 2022technical
In recent years, joint text-image embeddings have significantly improved thanks to the development of transformer-based Vision-Language models. Despite these advances, we still need to better understand the representations produced by those models. In this paper, we compare pre-trained and fine-tune…