2023
Dynamic Inference With Grounding Based Vision and Language Models
CVPR 2023poster
Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks.…