← Search

Amanmeet Garg

2 accepted papers

2023

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

ICCV 2023oral

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality while disregarding the audio signal for this task. Nevertheless, a recent advance…

Cited by 20PDFScholar
2023

Dynamic Inference With Grounding Based Vision and Language Models

CVPR 2023poster

Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks.…