← Search

Piyush Singh Pasi

2 accepted papers

2024

WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models

NeurIPS 2024poster

Cross-modal (image-to-text and text-to-image) retrieval is an established task used in evaluation benchmarks to test the performance of vision-language models (VLMs). Several state-of-the-art VLMs (e.g. CLIP, BLIP-2) have achieved near-perfect performance on widely-used image-text retrieval benchmar…

Cited by 0SourceScholar
2023

Temporally Aligning Long Audio Interviews with Questions: A Case Study in Multimodal Data Integration

IJCAI 2023poster

The problem of audio-to-text alignment has seen significant amount of research using complete supervision during training. However, this is typically not in the context of long audio recordings wherein the text being queried does not appear verbatim within the audio file. This work is a collaboratio…