← Search

Zhiqian Zhao

2 accepted papers

2026

Temporal Calibrating and Distilling for Scene-Text Aware Text-Video Retrieval

AAAI 2026technical

Existing text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing

Cited by 0SourcePDFScholar
2025

Heterogeneous Prompt-Guided Entity Inferring and Distilling for Scene-Text Aware Cross-Modal Retrieval

AAAI 2025technical

In cross-modal retrieval, comprehensive image understanding is vital while the scene text in images can provide fine-grained information to understand visual semantics. Current methods fail to make full use of scene text. They suffer from the semantic ambiguity of independent scene text and overlook…

Cited by 0SourcePDFScholar