2025
X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
EMNLP 2025
Prevalent text-to-video retrieval systems mainly adopt embedding models for feature extraction and compute cosine similarities for ranking. However, this design presents two limitations. Low-quality text-video data pairs could compromise the retrieval, yet are hard to identify and examine. Cosine si