2023
When and Why Vision-Language Models Behave like Bags-Of-Words, and What to Do About It?
ICLR 2023top-5%
Despite the success of large vision and language models (VLMs) in many downstream applications, it is unclear how well they encode the compositional relationships between objects and attributes. Here, we create the Attribution, Relation, and Order (ARO) benchmark to systematically evaluate the abili…