← Search

Dayan Wu

21 accepted papers

2026

Absorbing Gradient Conflicts: Modeling Semantic Variance via Kent Distributions for Cross-Modal Hashing

IJCAI 2026

Supervised proxy-based deep cross-modal hashing has become the dominant paradigm for large-scale retrieval. However, prevalent methods model class proxies as deterministic points in the embedding space. This rigid assumption causes severe gradient conflicts in multi-label scenarios, where gradient c

Cited by 0Scholar
2026

Discretization Is Not Always Better: Rethinking Deep Quantization for Asymmetric Image Retrieval

AAAI 2026technical

Asymmetric image retrieval (AIR), which typically employs a compact model for the query side and a large model for the database server, has garnered significant attention in resource-constrained environments. While deep hashing methods have shown great potential in large-scale image retrieval, curre

Cited by 0SourcePDFScholar
2026

EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval

CVPR 2026

Text-video retrieval tasks have seen significant improvements due to the recent development of large-scale vision-language pre-trained models. Traditional methods primarily focus on video representations or cross-modal alignment, while recent works shift toward enriching text expressiveness to bette

Cited by 0SourcecodeScholar
2026

Online Self-Calibration Against Hallucination in Vision-Language Models

IJCAI 2026

Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment methods typically rely on supervision distilled from stronger models such as GPT. However, this offline paradigm introdu

Cited by 0Scholar
2025

Categorical Attention: Fine-grained Language-guided Noise Filtering Network for Occluded Person Re-Identification

IJCAI 2025

Person Re-Identification (ReID) aims to match individuals across different camera views, but occlusions in real-world scenarios, such as vehicles or crowds, hinder feature extraction and matching. Current occluded ReID methodologies typically leverage visual augmentation techniques in an attempt to

Cited by 0SourcePDFScholar
2025

Endogenous Recovery via Within-modality Prototypes for Incomplete Multimodal Hashing

IJCAI 2025

Multimodal hashing projects multimodal data into compact binary codes, enabling rapid and storage-efficient retrieval of large-scale multimedia content. In practical scenarios, the issue of missing modality frequently arises when dealing with multimodal data. Existing incomplete multimodal hashing t

2024

A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re-Identification

CVPR 2024highlight

Extensive advancements have been made in person ReID through the mining of semantic information. Nevertheless existing methods that utilize semantic-parts from a single image modality do not explicitly achieve this goal. Whiteness the impressive capabilities in multimodal understanding of Vision Lan…

Cited by 12SourcePDFScholar
2024

Exploring Targeted Universal Adversarial Attack for Deep Hashing

ICASSP 2024accepted

Although image-dependent adversarial attacks have been studied, the more challenging image-agnostic adversarial attack for deep hashing remains an unexplored territory. In this paper, we take the first attempt on the more efficient and malicious targeted universal adversarial attack (TUAA) for deep…

Cited by 0SourceScholar
2024

Pairwise-Label-Based Deep Incremental Hashing with Simultaneous Code Expansion

AAAI 2024technical

Deep incremental hashing has become a subject of considerable interest due to its capability to learn hash codes in an incremental manner, eliminating the need to generate codes for classes that have already been learned. However, accommodating more classes requires longer hash codes, and regenerati…

Cited by 6SourcePDFScholar
2024

Prediction Exposes Your Face: Black-box Model Inversion via Prediction Alignment

ECCV 2024poster

"Model inversion (MI) attack reconstructs the private training data of a target model given its output, posing a significant threat to deep learning models and data privacy. On one hand, most of existing MI methods focus on searching for latent codes to represent the target identity, yet this iterat…

2023

AREA: Adaptive Reweighting via Effective Area for Long-Tailed Classification

ICCV 2023poster

Large-scale data from real-world usually follow a long-tailed distribution (i.e., a few majority classes occupy plentiful training data, while most minority classes have few samples), making the hyperplanes heavily skewed to the minority classes. Traditionally, reweighting is adopted to make the hyp…

Cited by 47PDFcodeScholar
2023

Dual Alignment Unsupervised Domain Adaptation for Video-Text Retrieval

CVPR 2023poster

Video-text retrieval is an emerging stream in both computer vision and natural language processing communities, which aims to find relevant videos given text queries. In this paper, we study the notoriously challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein…

Cited by 24SourcePDFScholar
2023

TeAw: Text-Aware Few-Shot Remote Sensing Image Scene Classification

ICASSP 2023accepted

The recent advance has shown that few-shot learning may be a promising way to alleviate the data reliance of remote sensing image scene classification. However, most existing works focus on extracting distinguishable features only from visual modality, while the problem of learning knowledge from mu…

Cited by 0SourceScholar
2022

Clustering and Separating Similarities for Deep Unsupervised Hashing

ICASSP 2022accepted

The lack of supervised information is the pivotal problem in unsupervised hashing. Most methods leverage deep features extracted from pre-trained models to generate semantic similarities as supervised information. These fixed features are, however, neither designed originally for retrieval nor updat…

Cited by 0SourceScholar
2022

Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed Classification

AAAI 2022technical

Real-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles…

2022

Prototype-Based Inter-Camera Learning for Person Re-Identification

ICASSP 2022accepted

Person re-identification (ReID) aims at retrieving images of the same person across non-overlapping camera views. The prior works focus on either fully supervised or unsupervised ReID settings, and achieve remarkable performances. In real scenarios, however, the major annotation cost comes from matc…

Cited by 0SourceScholar
2021

FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text Detection

ICASSP 2021accepted

Accurate detection of multi-oriented text that accounts for a large proportion in real practice is of great significance. The performance has improved rapidly on common benchmarks in recent years. However, dense long text case and the quality of detection are easy to be overlooked. Direct regression…

Cited by 0SourceScholar
2018

Deep Uniqueness-Aware Hashing for Fine-Grained Multi-Label Image Retrieval

ICASSP 2018accepted

Deep supervised hashing methods for multi-label image retrieval have achieved great success nowadays. However, these methods only take the similarity between the database images and the query images into account, but they ignore the uniqueness of the database images when deciding on their rankings.…

Cited by 0SourceScholar