← Search

Jingru Yi

1 accepted papers

2023

Dynamic Inference With Grounding Based Vision and Language Models

CVPR 2023poster

Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks.…