← Search

Howoong Lee

2 accepted papers

2025

Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization

EMNLP 2025

In text-video retrieval, auxiliary captions are often used to enhance video understanding, bridging the gap between the modalities. While recent advances in multi-modal large language models (MLLMs) have enabled strong zero-shot caption generation, we observe that such captions tend to be generic an