← Search

Wanqian Zhang

10 accepted papers

2026

Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

CVPR 2026

Open-Vocabulary Temporal Action Detection (OV-TAD) aims to classify and localize action segments in untrimmed videos for unseen categories. Previous methods rely solely on global alignment between label-level semantics and visual features, which is insufficient to transfer temporal consistent visual

Cited by 0SourceScholar
2025

AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts

ICCV 2025poster

Despite rapid advancements in text-to-image (T2I) models, their safety mechanisms are vulnerable to adversarial prompts, which maliciously generate unsafe images. Current red-teaming methods for proactively assessing such vulnerabilities usually require white-box access to T2I models, and rely on in…

Cited by 0SourcePDFScholar
2025

CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR

CVPR 2025poster

Scene Text Retrieval (STR) seeks to identify all images containing a given query string. Existing methods typically rely on an explicit Optical Character Recognition (OCR) process of text spotting or localization, which is susceptible to complex pipelines and accumulated errors. To settle this, we r…

Cited by 0SourcePDFScholar
2024

Exploring Targeted Universal Adversarial Attack for Deep Hashing

ICASSP 2024accepted

Although image-dependent adversarial attacks have been studied, the more challenging image-agnostic adversarial attack for deep hashing remains an unexplored territory. In this paper, we take the first attempt on the more efficient and malicious targeted universal adversarial attack (TUAA) for deep…

Cited by 0SourceScholar
2024

Prediction Exposes Your Face: Black-box Model Inversion via Prediction Alignment

ECCV 2024poster

"Model inversion (MI) attack reconstructs the private training data of a target model given its output, posing a significant threat to deep learning models and data privacy. On one hand, most of existing MI methods focus on searching for latent codes to represent the target identity, yet this iterat…

2023

Dual Alignment Unsupervised Domain Adaptation for Video-Text Retrieval

CVPR 2023poster

Video-text retrieval is an emerging stream in both computer vision and natural language processing communities, which aims to find relevant videos given text queries. In this paper, we study the notoriously challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein…

Cited by 24SourcePDFScholar
2022

Clustering and Separating Similarities for Deep Unsupervised Hashing

ICASSP 2022accepted

The lack of supervised information is the pivotal problem in unsupervised hashing. Most methods leverage deep features extracted from pre-trained models to generate semantic similarities as supervised information. These fixed features are, however, neither designed originally for retrieval nor updat…

Cited by 0SourceScholar
2022

Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

AAAI 2022technical

Real-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles…

2022

Prototype-Based Inter-Camera Learning for Person Re-Identification

ICASSP 2022accepted

Person re-identification (ReID) aims at retrieving images of the same person across non-overlapping camera views. The prior works focus on either fully supervised or unsupervised ReID settings, and achieve remarkable performances. In real scenarios, however, the major annotation cost comes from matc…

Cited by 0SourceScholar