Diffusion Models are Zero-Shot Generative Text-Vision Retrievers
Bao Li, Zeke Xie, Xiaomei Zhang, Xiangyu Zhu, Zhen Lei
Abstract
Large-scale text-to-image diffusion models have demonstrated impressive capabilities for downstream tasks by leveraging strong vision-language alignment from generative pre-training. Recently, a number of works have explored how to use the power of text-to-image diffusion models for text-image matching tasks. While previous generative text-image matching methods have shown potential in retrieving the most challenging candidates, they still suffer from extremely slow retrieval speeds and lack the ability to handle temporal dimensions, making them impractical for text-video retrieval. In this paper, we propose ZSGenRet, a simple yet effective zero-shot generative text-vision retrieval framework for both text-image and text-video tasks, based on pre-trained text-to-image diffusion models. We further incorporate inversion saliency detection to identify key frames in videos and enhance the semantic representation of the vision encoder. Experimental results demonstrate that ZSGenRet significantly improves text-video retrieval performance and achieves competitive results on text-image retrieval while remarkably improving efficiency. To the best of our knowledge, the proposed ZSGenRet is the first to explore zero-shot text-video retrieval based on diffusion models.
BibTeX
@inproceedings{icassp2025_diffusionmodelsa,
title = {Diffusion Models are Zero-Shot Generative Text-Vision Retrievers},
author = {Bao Li and Zeke Xie and Xiaomei Zhang and Xiangyu Zhu and Zhen Lei},
booktitle = {ICASSP 2025},
year = {2025}
}