Do Vision-and-Language Transformers Learn Grounded Predicate-Noun Dependencies?
Recent advances in vision-and-language modeling have seen the development of Transformer architectures that achieve remarkable performance on multimodal reasoning tasks.Yet, the exact capabilities of these black-box models are still poorly understood. While much of previous work has focused on study…