2023
CREPE: Can Vision-Language Foundation Models Reason Compositionally?
CVPR 2023highlight
A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, we find that--across 7 architectures trained with 4 algorithms on massive datasets--they struggle at c…