← Search

Emmanuelle Salin

2 accepted papers

2022

Are Vision-Language Transformers Learning Multimodal Representations? A Probing Perspective

AAAI 2022technical

In recent years, joint text-image embeddings have significantly improved thanks to the development of transformer-based Vision-Language models. Despite these advances, we still need to better understand the representations produced by those models. In this paper, we compare pre-trained and fine-tune…

2022

Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?

EMNLP 2022main

Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks.Yet, the exact capabilities of these black-box models are still poorly understood. While much of previous work has focused on study…