← Search

Zequn Xie

4 accepted papers

2026

CONQUER: CONTEXT-AWARE REPRESENTATION WITH QUERY ENHANCEMENT FOR TEXT-BASED PERSON SEARCH

ICASSP 2026poster

Text-Based Person Search (TBPS) aims to retrieve pedestrian images from large galleries using natural language descriptions. This task, essential for public safety applications, is hindered by cross-modal discrepancies and ambiguous user queries. We introduce CONQUER, a two-stage framework designed…

Cited by 0SourcePDFScholar
2026

HVD: HUMAN VISION-DRIVEN VIDEO REPRESENTATION LEARNING FOR TEXT-VIDEO RETRIEVAL

ICASSP 2026poster

The success of CLIP has driven substantial progress in text-video retrieval. However, current methods often suffer from "blind" feature interaction, where the model struggles to discern key visual information from background noise due to the sparsity of textual queries. To bridge this gap, we draw i…

Cited by 0SourcePDFScholar
2026

Scene-Aware Spatiotemporal Generalization: Towards Robust Temporal Action Detection Across Domains

AAAI 2026technical

Temporal Action Detection (TAD) aims to identify specific actions in long, untrimmed videos by determining their start, end times and categories, yet existing models suffer from performance degradation under out-of-distribution scenarios due to unrealistic i.i.d. assumptions. While domain generaliza

Cited by 0SourcePDFScholar
2025

Chat-Driven Text Generation and Interaction for Person Retrieval

EMNLP 2025

Text-based person search (TBPS) enables the retrieval of person images from large-scale databases using natural language descriptions, offering critical value in surveillance applications. However, a major challenge lies in the labor-intensive process of obtaining high-quality textual annotations, w

Cited by 0SourcePDFScholar