← Search

Yongjian Wu

34 accepted papers

2025

Visual Textualization for Image Prompted Object Detection

ICCV 2025poster

We propose VisTex-OVLM, a novel image prompted object detection method that introduces visual textualization ---- a process that projects a few visual exemplars into the text feature space to enhance Object-level Vision-Language Models' (OVLMs) capability in detecting rare categories that are diffic…

Cited by 0SourcePDFScholar
2024

Attention Disturbance and Dual-Path Constraint Network for Occluded Person Re-identification

AAAI 2024technical

Occluded person re-identification (Re-ID) aims to address the potential occlusion problem when matching occluded or holistic pedestrians from different camera views. Many methods use the background as artificial occlusion and rely on attention networks to exclude noisy interference. However, the si…

Cited by 10SourcePDFScholar
2024

Occluded Person Re-identification via Saliency-Guided Patch Transfer

AAAI 2024technical

While generic person re-identification has made remarkable improvement in recent years, these methods are designed under the assumption that the entire body of the person is available. This assumption brings about a significant performance degradation when suffering from occlusion caused by various…

Cited by 16SourcePDFScholar
2024

RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification

NeurIPS 2024poster

This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this…

2024

SDPT: Synchronous Dual Prompt Tuning for Fusion-based Visual-Language Pre-trained Models

ECCV 2024poster

"Prompt tuning methods have achieved remarkable success in parameter-efficient fine-tuning on large pre-trained models. However, their application to dual-modal fusion-based visual-language pre-trained models (VLPMs), such as GLIP, has encountered issues. Existing prompt tuning methods have not effe…

2023

CF-ViT: A General Coarse-to-Fine Method for Vision Transformer

AAAI 2023technical

Vision Transformers (ViT) have made many breakthroughs in computer vision tasks. However, considerable redundancy arises in the spatial dimension of an input image, leading to massive computational costs. Therefore, We propose a coarse-to-fine vision transformer (CF-ViT) to relieve computational bur…

2023

Improving Adversarial Robustness via Information Bottleneck Distillation

NeurIPS 2023poster

Previous studies have shown that optimizing the information bottleneck can significantly improve the robustness of deep neural networks. Our study closely examines the information bottleneck principle and proposes an Information Bottleneck Distillation approach. This specially designed, robust disti…

2023

OMPQ: Orthogonal Mixed Precision Quantization

AAAI 2023technical

To bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash…

2023

Towards Real-Time Panoptic Narrative Grounding by an End-to-End Grounding Network

AAAI 2023technical

Panoptic Narrative Grounding (PNG) is an emerging cross-modal grounding task, which locates the target regions of an image corresponding to the text description. Existing approaches for PNG are mainly based on a two-stage paradigm, which is computationally expensive. In this paper, we propose a one-…

2022

Black-Box Dissector: Towards Erasing-Based Hard-Label Model Stealing Attack

ECCV 2022poster

"Previous studies have verified that the functionality of black-box models can be stolen with full probability outputs. However, under the more practical hard-label setting, we observe that existing methods suffer from catastrophic performance degradation. We argue this is due to the lack of rich in…

2022

Dynamic Dual Trainable Bounds for Ultra-Low Precision Super-Resolution Networks

ECCV 2022poster

"Light-weight super-resolution (SR) models have received considerable attention for their serviceability in mobile devices. Many efforts employ network quantization to compress SR models. However, these methods suffer from severe performance degradation when quantizing the SR models to ultra-low pre…

2022

Fine-Grained Data Distribution Alignment for Post-Training Quantization

ECCV 2022poster

"While post-training quantization receives popularity mostly due to its evasion in accessing the original complete training dataset, its poor performance also stems from scarce images. To alleviate this limitation, in this paper, we leverage the synthetic data introduced by zero-shot quantization wi…

2022

Learning Best Combination for Efficient N:M Sparsity

NeurIPS 2022accept

By forcing N out of M consecutive weights to be non-zero, the recent N:M fine-grained network sparsity has received increasing attention with its two attractive advantages over traditional irregular network sparsity methods: 1) Promising performance at a high sparsity. 2) Significant speedups when p…

2021

Aha! Adaptive History-Driven Attack for Decision-Based Black-Box Models

ICCV 2021poster

The decision-based black-box attack means to craft adversarial examples with only the top-1 label of the victim model available. A common practice is to start from a large perturbation and then iteratively reduce it with a deterministic direction and a random one while keeping it adversarial. The li…

Cited by 21PDFScholar
2021

Discover Cross-Modality Nuances for Visible-Infrared Person Re-Identification

CVPR 2021poster

Visible-infrared person re-identification (Re-ID) aims to match the pedestrian images of the same identity from different modalities. Existing works mainly focus on alleviating the modality discrepancy by aligning the distributions of features from different modalities. However, nuanced but discrimi…

Cited by 286PDFcodeScholar
2021

Dual-level Collaborative Transformer for Image Captioning

AAAI 2021technical

Descriptive region features extracted by object detection networks have played an important role in the recent advancements of image captioning. However, they are still criticized for the lack of contextual information and fine-grained details, which in contrast are the merits of traditional grid fe…

2021

HifiFace: 3D Shape and Semantic Prior Guided High Fidelity Face Swapping

IJCAI 2021poster

In this work, we propose a high fidelity face swapping method, called HifiFace, which can well preserve the face shape of the source face and generate photo-realistic results. Unlike other existing face swapping works that only use face recognition model to keep the identity similarity, we propose 3…

2021

Image-to-Image Translation via Hierarchical Style Disentanglement

CVPR 2021poster

Recently, image-to-image translation has made significant progress in achieving both multi-label (i.e., translation conditioned on different labels) and multi-style (i.e., generation with diverse styles) tasks. However, due to the unexplored independence and exclusiveness in the labels, existing end…

Cited by 159PDFcodeScholar
2021

Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network

AAAI 2021technical

Transformer-based architectures have shown great success in image captioning, where object regions are encoded and then attended into the vectorial representations to guide the caption decoding. However, such vectorial representations only contain region-level information without considering the glo…

Cited by 208SourcePDFScholar
2021

Parallel Detection-and-Segmentation Learning for Weakly Supervised Instance Segmentation

ICCV 2021poster

Weakly supervised instance segmentation (WSIS) with only image-level labels has recently drawn much attention. To date, bottom-up WSIS methods refine discriminative cues from classifiers with sophisticated multi-stage training procedures, which also suffer from inconsistent object boundaries. And to…

Cited by 23PDFScholar
2021

RSTNet: Captioning With Adaptive Attention on Visual and Non-Visual Words

CVPR 2021poster

Recent progress on visual question answering has explored the merits of grid features for vision language tasks. Meanwhile, transformer-based models have shown remarkable performance in various sequence prediction problems. However, the spatial information loss of grid features caused by flattening…

Cited by 286PDFcodeScholar
2021

Toward Joint Thing-and-Stuff Mining for Weakly Supervised Panoptic Segmentation

CVPR 2021poster

Panoptic segmentation aims to partition an image to object instances and semantic content for thing and stuff categories, respectively. To date, learning weakly supervised panoptic segmentation (WSPS) with only image-level labels remains unexplored. In this paper, we propose an efficient jointly thi…

Cited by 19PDFScholar
2020

Channel Pruning via Automatic Structure Search

IJCAI 2020poster

Channel pruning is among the predominant approaches to compress deep neural networks. To this end, most existing pruning methods focus on selecting channels (filters) by importance/optimization or regularization based on rule-of-thumb designs, which defects in sub-optimal pruning. In this paper, we…

2020

Interpretable Neural Network Decoupling

ECCV 2020poster

The remarkable performance of convolutional neural networks (CNNs) is entangled with their huge number of uninterpretable parameters, which has become the bottleneck limiting the exploitation of their full potential. Towards network interpretation, previous endeavors mainly resort to the single filt…

Cited by 10SourcePDFScholar
2020

Rotated Binary Neural Network

NeurIPS 2020poster

Binary Neural Network (BNN) shows its predominance in reducing the complexity of deep neural networks. However, it suffers severe performance degradation. One of the major impediments is the large quantization error between the full-precision weight vector and its binary vector. Previous works focus…

2020

UWSOD: Toward Fully-Supervised-Level Capacity Weakly Supervised Object Detection

NeurIPS 2020poster

Weakly supervised object detection (WSOD) has attracted extensive research attention due to its great flexibility of exploiting large-scale dataset with only image-level annotations for detector training. Despite its great advance in recent years, WSOD still suffers limited performance, which is far…

2019

Cyclic Guidance for Weakly Supervised Joint Detection and Segmentation

CVPR 2019poster

Weakly supervised learning has attracted growing research attention due to the significant saving in annotation cost for tasks that require intra-image annotations, such as object detection and semantic segmentation. To this end, existing weakly supervised object detection and semantic segmentation…

Cited by 142PDFcodeScholar
2019

Exploiting Kernel Sparsity and Entropy for Interpretable CNN Compression

CVPR 2019poster

Compressing convolutional neural networks (CNNs) has received ever-increasing research focus. However, most existing CNN compression methods do not interpret their inherent structures to distinguish the implicit redundancy. In this paper, we investigate the problem of CNN compression from a novel in…

Cited by 177PDFcodeScholar
2019

Universal Adversarial Perturbation via Prior Driven Uncertainty Approximation

ICCV 2019oral

Deep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupe…

Cited by 119PDFScholar
2019

Variational Structured Semantic Inference for Diverse Image Captioning

NeurIPS 2019poster

Despite the exciting progress in image captioning, generating diverse captions for a given image remains as an open problem. Existing methods typically apply generative models such as Variational Auto-Encoder to diversify the captions, which however neglect two key factors of diverse expression, i.e…

2019

Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch Refinement

ICASSP 2019accepted

Data-driven methods for melody extraction from polyphonic music generally require large amounts of labeled data for model training. However, musical data with annotations of melody fundamental frequency (F0) are rare and hard to obtain. To overcome this limitation, in this paper we propose to use me…

Cited by 0SourceScholar
2018

GroupCap: Group-Based Image Captioning With Structured Relevance and Diversity Constraints

CVPR 2018poster

Most image captioning models focus on one-line (single image) captioning, where the correlations like relevance and diversity among group images (e.g., within the same album or event) are simply neglected, resulting in less accurate and diverse captions. Recent works mainly consider imposing the div…

2017

Cross-Modality Binary Code Learning via Fusion Similarity Hashing

CVPR 2017poster

Binary code learning has been emerging topic in large-scale cross-modality retrieval recently. It aims to map features from multiple modalities into a common Hamming space, where the cross-modality similarity can be approximated efficiently via Hamming distance. To this end, most existing works lear…

Cited by 255PDFScholar
2017

Fusing transcription results from polyphonic and monophonic audio for singing melody transcription in polyphonic music

ICASSP 2017accepted

This paper presents a new system for singing melody transcription from polyphonic songs. Instead of operating solely on polyphonic audio of each song to be processed (as most existing systems do), our system takes as inputs additionally multiple monophonic recordings of people singing the song. To t…

Cited by 0SourceScholar