2023
Incorporating Structured Representations into Pretrained Vision \& Language Models Using Scene Graphs
EMNLP 2023long main
Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks. However, recent works have shown that even the best VLMs struggle to capture aspects of compositional scene understanding, such as object attributes, relations, and action states. In cont…