← Search

Nannan Wang

77 accepted papers

2026

Harnessing Textual Semantic Priors for Knowledge Transfer and Refinement in CLIP-Driven Continual Learning

AAAI 2026technical

Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-training (CLIP), their promise for CL has attracted increasing attention due to their s

Cited by 0SourcePDFScholar
2026

Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

CVPR 2026

Incremental Object Detection (IOD) enables AI systems to continuously acquire new object classes while preserving knowledge of previously learned ones, an ability essential for deployment in dynamic, real-world environments. Existing IOD methods typically rely on knowledge distillation to mitigate c

Cited by 0SourceScholar
2026

Interference-Isolated Elastic Weight Consolidation and Knowledge Calibration for Incremental Object Detection

ICLR 2026poster

Incremental Object Detection (IOD) enables AI systems to continuously learn new object classes over time while retaining knowledge of previously learned categories. This capability is essential for adapting to dynamic environments without forgetting prior information. Although existing IOD methods h…

Cited by 0SourceScholar
2026

Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution

AAAI 2026technical

The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-t

Cited by 0SourcePDFScholar
2026

PSMix: Robust Point Cloud Recognition through Spectral Domain Mixing

ICML 2026poster

While data augmentation is essential for robust point cloud recognition, conventional spatial mixup strategies often compromise geometric integrity by generating physically unrealistic samples. To overcome this limitation, we propose PSMix, which shifts the mixing paradigm to the spectral domain via…

Cited by 0SourceScholar
2026

Reasoning-Driven Multimodal LLM for Domain Generalization

ICLR 2026poster

This paper addresses the domain generalization (DG) problem in deep learning. While most DG methods focus on enforcing visual feature invariance, we leverage the reasoning capability of multimodal large language models (MLLMs) and explore the potential of constructing reasoning chains that derives…

Cited by 0SourceScholar
2026

Revealing the Invisible: Latent Structure Modeling for Semantically Consistent Cloud Removal

AAAI 2026technical

Cloud removal (CR) in remote sensing imagery is a critical yet challenging task due to complex cloud patterns and diverse underlying ground structures. Despite recent progress in generative models such as diffusion models, CR remains limited by their inadequate capability to perceive and reconstruct

Cited by 0SourcePDFScholar
2026

Robust Cross-Modal Retrieval via Generative Semantic Refinement and Exclusion-Guided Adaptation

ICML 2026poster

Vision-Language Pre-trained (VLP) models are vulnerable to real-world query noise. Current cross-modal Test-Time Adaptation (TTA) methods often rely on high-confidence predictions, which induces confirmation bias and neglects the informative signals in ambiguous Low-Confidence Queries. To address th…

Cited by 0SourceScholar
2026

SPR$^2$Q: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution

ICLR 2026poster

Low-bit quantization has achieved significant progress in image super-resolution. However, existing quantization methods show evident limitations in handling the heterogeneity of different components. Particularly under extreme low-bit compression, the issue of information loss becomes especially pr…

Cited by 0SourceScholar
2026

StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning

ICLR 2026poster

Video Class-Incremental Learning (VCIL) seeks to develop models that continuously learn new action categories over time without forgetting previously acquired knowledge. Unlike traditional Class-Incremental Learning (CIL), VCIL introduces the added complexity of spatiotemporal structures, making it…

Cited by 0SourceScholar
2026

Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection

ICML 2026poster

Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learning perspective, separating feature spaces to sharpen decision boundaries. However, this paradigm conflicts with the inhe…

Cited by 0SourceScholar
2026

Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning

ICML 2026poster

Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low-Rank Adaptation (LoRA), a representative Parameter-Efficient Fine-Tuning (PEFT) method, has gained increasing attention in CL. Several LoRA-based CL methods reduce interference …

Cited by 0SourceScholar
2026

When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondence

CVPR 2026

Establishing accurate correspondence between sparse line representations and rich textured imagery remains a formidable challenge. While diffusion features excel in semantic correspondence, they struggle to bridge the fundamental gap between abstract sketches and texture-rich photographs. We identif

Cited by 0SourcecodeScholar
2025

3D Test-time Adaptation via Graph Spectral Driven Point Shift

ICCV 2025poster

While test-time adaptation (TTA) methods effectively address domain shifts by dynamically adapting pre-trained models to target domain data during online inference, their application to 3D point clouds is hindered by their irregular and unordered structure. Current 3D TTA methods often rely on compu…

Cited by 0SourcePDFScholar
2025

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

CVPR 2025poster

Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data t…

Cited by 0SourcePDFScholar
2025

Asymmetric Reinforcing Against Multi-Modal Representation Bias

AAAI 2025technical

The strength of multimodal learning lies in its ability to integrate information from various sources, providing rich and comprehensive insights. However, in real-world scenarios, multi-modal systems often face the challenge of dynamic modality contributions, the dominance of different modalities ma…

2025

DIH-CLIP: Unleashing the Diversity of Multi-Head Self-Attention for Training-Free Open-Vocabulary Semantic Segmentation

ICCV 2025poster

Recent Training-Free Open-Vocabulary Semantic Segmentation (TF-OVSS) leverages a pre-training vision-language model to segment images from open-set visual concepts without training and fine-tuning. The key of TF-OVSS is to improve the local spatial representation of CLIP by leveraging self-correlati…

2025

Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts

ICML 2025poster

Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformer…

Cited by 0SourcePDFScholar
2025

Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object Detection

ICCV 2025poster

Domain incremental object detection in remote sensing addresses the challenge of adapting to continuously emerging domains with distinct characteristics. Unlike natural images, remote sensing data vary significantly due to differences in sensors, altitudes, and geographic locations, leading to data…

Cited by 0SourcePDFScholar
2025

Effective Diffusion Transformer Architecture for Image Super-Resolution

AAAI 2025technical

Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image…

2025

Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement

AAAI 2025technical

Adversarial fine-tuning methods enhance adversarial robustness via fine-tuning the pre-trained model in an adversarial training manner. However, we identify that some specific latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap b…

Cited by 0SourcePDFScholar
2025

Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model

AAAI 2025technical

Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (f…

2025

Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation

CVPR 2025highlight

Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate class-specific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classifica…

Cited by 0SourcePDFScholar
2025

Phase and Amplitude-aware Prompting for Enhancing Adversarial Robustness

ICML 2025poster

Deep neural networks are found to be vulnerable to adversarial perturbations. The prompt-based defense has been increasingly studied due to its high efficiency. However, existing prompt-based defenses mainly exploited mixed prompt patterns, where critical patterns closely related to object semantics…

Cited by 0SourcePDFScholar
2025

Q-Norm: Robust Representation Learning via Quality-Adaptive Normalization

ICCV 2025poster

Although deep neural networks have achieved remarkable success in various computer vision tasks, they face significant challenges in degraded image understanding due to domain shifts caused by quality variations. Drawing biological inspiration from the human visual system (HVS), which dynamically ad…

2025

QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception

AAAI 2025technical

Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To…

2025

ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence Learning

CVPR 2025poster

Can we accurately identify the true correspondences from multimodal datasets containing mismatched data pairs? Existing methods primarily emphasize the similarity matching between the representations of objects across modalities, potentially neglecting the crucial relation consistency within modalit…

2025

Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language Models

ICML 2025poster

Prompt learning is a cutting-edge parameter-efficient fine-tuning technique for pre-trained vision-language models (VLMs). Instead of learning a single text prompt, recent works have revealed that learning diverse text prompts can effectively boost the performances on downstream tasks, as the divers…

Cited by 0SourcePDFScholar
2025

Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations

AAAI 2025technical

Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fai…

2025

Towards Regularized Mixture of Predictions for Class-Imbalanced Semi-Supervised Facial Expression Recognition

IJCAI 2025

Semi-supervised facial expression recognition (SSFER) effectively assigns pseudo-labels to confident unlabeled samples when only limited emotional annotations are available. Existing SSFER methods are typically built upon an assumption of the class-balanced distribution. However, they are far from r

2025

Training Consistent Mixture-of-Experts-Based Prompt Generator for Continual Learning

AAAI 2025technical

Visual prompt tuning-based continual learning (CL) methods have shown promising performance in exemplar-free scenarios, where their key component can be viewed as a prompt generator. Existing approaches generally rely on freezing old prompts, slow updating and task discrimination for prompt generato…

Cited by 0SourcePDFScholar
2024

Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model

AAAI 2024technical

Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low t…

2024

Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors

IJCAI 2024poster

The remarkable prowess of diffusion models in image generation has spurred efforts to extend their application beyond generative tasks. However, a persistent challenge exists in lacking a unified approach to apply diffusion models to visual perception tasks with diverse semantic granularity requirem…

Cited by 3SourcePDFScholar
2024

Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection

NeurIPS 2024poster

Unsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for developing a safe real-world machine learning system. Current reconstruction-based method provides a good alternative appro…

2024

Disentangled Prompt Representation for Domain Generalization

CVPR 2024poster

Domain Generalization (DG) aims to develop a versatile model capable of performing well on unseen target domains. Recent advancements in pre-trained Visual Foundation Models (VFMs) such as CLIP show significant potential in enhancing the generalization abilities of deep models. Although there is a g…

Cited by 9SourcePDFScholar
2024

Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification

NeurIPS 2024poster

Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of mode…

Cited by 1SourcePDFScholar
2024

Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End Pipeline

CVPR 2024poster

In this paper we explore a novel challenging generation task i.e. Handwritten Mathematical Expression Generation (HMEG) from symbolic sequences. Since symbolic sequences are naturally graph-structured data we formulate HMEG as a graph-to-image (G2I) generation problem. Unlike the generation of natur…

2024

Human-Robot Interactive Creation of Artistic Portrait Drawings

ICRA 2024poster

In this paper, we present a novel system for Human-Robot Interactive Creation of Artworks (HRICA). Different from previous robot painters, HRICA allows a human user and a robot to alternately draw strokes on a canvas, to collaboratively create a portrait drawing through frequent interactions. The ke…

Cited by 0SourcecodeScholar
2024

Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search

IJCAI 2024poster

One-step Weakly Supervised Person Search (WSPS) jointly performs pedestrian detection and person Re-IDentification (ReID) only with bounding box annotations, which makes the traditional person ReID problem more suitable and efficient for real-world applications. However, this task is very challengin…

Cited by 0SourcePDFScholar
2024

Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking

AAAI 2024technical

The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature pos…

Cited by 4SourcePDFScholar
2024

On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection

ICLR 2024poster

Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection d…

Cited by 2SourcePDFScholar
2024

Point Deformable Network with Enhanced Normal Embedding for Point Cloud Analysis

AAAI 2024technical

Recently MLP-based methods have shown strong performance in point cloud analysis. Simple MLP architectures are able to learn geometric features in local point groups yet fail to model long-range dependencies directly. In this paper, we propose Point Deformable Network (PDNet), a concise MLP-based ne…

Cited by 3SourcePDFScholar
2024

Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization

ECCV 2024poster

"Weakly Supervised Object Localization (WSOL), which aims to localize objects by only using image-level labels, has attracted much attention because of its low annotation cost in real applications. Current studies focus on the Class Activation Map (CAM) of CNN and the self-attention map of transform…

Cited by 1SourcePDFScholar
2024

Robust Training of Federated Models with Extremely Label Deficiency

ICLR 2024poster

Federated semi-supervised learning (FSSL) has emerged as a powerful paradigm for collaboratively training machine learning models using distributed data with label deficiency. Advanced FSSL methods predominantly focus on training a single model on each client. However, this approach could lead to a…

Cited by 8SourcePDFScholar
2024

SHaRPose: Sparse High-Resolution Representation for Human Pose Estimation

AAAI 2024technical

High-resolution representation is essential for achieving good performance in human pose estimation models. To obtain such features, existing works utilize high-resolution input images or fine-grained image tokens. However, this dense high-resolution representation brings a significant computational…

2024

Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning

ICML 2024poster

Continual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis---one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task--…

Cited by 7SourcePDFScholar
2024

Visual Prompt Tuning in Null Space for Continual Learning

NeurIPS 2024poster

Existing prompt-tuning methods have demonstrated impressive performances in continual learning (CL), by selecting and updating relevant prompts in the vision-transformer models. On the contrary, this paper aims to learn each task by tuning the prompts in the direction orthogonal to the subspace span…

2023

Boosting Weakly-Supervised Temporal Action Localization With Text Information

CVPR 2023poster

Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost WTAL from two aspects, i.e., (a) the discriminative objecti…

2023

Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding

AAAI 2023technical

Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which…

Cited by 14SourcePDFScholar
2023

Eliminating Adversarial Noise via Information Discard and Robust Representation Restoration

ICML 2023poster

Deep neural networks (DNNs) are vulnerable to adversarial noise. Denoising model-based defense is a major protection strategy. However, denoising models may fail and induce negative effects in fully white-box scenarios. In this work, we start from the latent inherent properties of adversarial sample…

Cited by 8SourcePDFScholar
2023

Hiding Visual Information via Obfuscating Adversarial Perturbations

ICCV 2023poster

Growing leakage and misuse of visual information raise security and privacy concerns, which promotes the development of information protection. Existing adversarial perturbations-based methods mainly focus on the de-identification against deep learning models. However, the inherent visual informatio…

Cited by 12PDFcodeScholar
2023

Masked and Adaptive Transformer for Exemplar Based Image Translation

CVPR 2023poster

We present a novel framework for exemplar based image translation. Recent advanced methods for this task mainly focus on establishing cross-domain semantic correspondence, which sequentially dominates image generation in the manner of local style control. Unfortunately, cross domain semantic matchin…

2023

NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning

NeurIPS 2023poster

As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and de…

2023

Phase-aware Adversarial Defense for Improving Adversarial Robustness

ICML 2023poster

Deep neural networks have been found to be vulnerable to adversarial noise. Recent works show that exploring the impact of adversarial noise on intrinsic components of data can help improve adversarial robustness. However, the pattern closely related to human perception has not been deeply studied.…

Cited by 8SourcePDFScholar
2023

Semantic-Aware Generation of Multi-View Portrait Drawings

IJCAI 2023poster

Neural radiance fields (NeRF) based methods have shown amazing performance in synthesizing 3D-consistent photographic images, but fail to generate multi-view portrait drawings. The key is that the basic assumption of these methods -- a surface point is consistent when rendered from different views -…

2022

Class-Dependent Label-Noise Learning with Cycle-Consistency Regularization

NeurIPS 2022accept

In label-noise learning, estimating the transition matrix plays an important role in building statistically consistent classifier. Current state-of-the-art consistent estimator for the transition matrix has been developed under the newly proposed sufficiently scattered assumption, through incorporat…

Cited by 40SourcePDFScholar
2022

Exploring Set Similarity for Dense Self-Supervised Representation Learning

CVPR 2022poster

By considering the spatial correspondence, dense self-supervised representation learning has achieved superior performance on various dense prediction tasks. However, the pixel-level correspondence tends to be noisy because of many similar misleading pixels, e.g., backgrounds. To address this issue,…

Cited by 51PDFcodeScholar
2022

Improving Adversarial Robustness via Mutual Information Estimation

ICML 2022spotlight

Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples f…

2022

Instance-Dependent Label-Noise Learning With Manifold-Regularized Transition Matrix Estimation

CVPR 2022poster

In label-noise learning, estimating the transition matrix has attracted more and more attention as the matrix plays an important role in building statistically consistent classifiers. However, it is very challenging to estimate the transition matrix T(x), where T(x) denotes the instance, because it…

Cited by 93PDFScholar
2022

Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction

IJCAI 2022poster

Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the res…

Cited by 7SourcePDFScholar
2022

Towards Semi-Supervised Deep Facial Expression Recognition With an Adaptive Confidence Margin

CVPR 2022poster

Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all…

Cited by 116PDFcodeScholar
2021

A Sketch-Transformer Network for Face Photo-Sketch Synthesis

IJCAI 2021poster

We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promisi…

Cited by 32SourcePDFScholar
2021

Class2Simi: A Noise Reduction Perspective on Learning with Noisy Labels

ICML 2021spotlight

Learning with noisy labels has attracted a lot of attention in recent years, where the mainstream approaches are in \emph{pointwise} manners. Meanwhile, \emph{pairwise} manners have shown great potential in supervised metric learning and unsupervised contrastive learning. Thus, a natural question is…

Cited by 82SourcePDFScholar
2021

Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer

CVPR 2021poster

Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize comple…

Cited by 119PDFcodeScholar
2021

Removing Adversarial Noise in Class Activation Feature Space

ICCV 2021poster

Deep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this…

Cited by 36PDFcodeScholar
2021

Robust early-learning: Hindering the memorization of noisy labels

ICLR 2021poster

The \textit{memorization effects} of deep networks show that they will first memorize training data with clean labels and then those with noisy labels. The \textit{early stopping} method therefore can be exploited for learning with noisy labels. However, the side effect brought by noisy labels will…

Cited by 354SourcePDFScholar
2021

Support-Set Based Cross-Supervision for Video Grounding

ICCV 2021poster

Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel…

Cited by 53PDFScholar
2021

Syncretic Modality Collaborative Learning for Visible Infrared Person Re-Identification

ICCV 2021poster

Visible infrared person re-identification (VI-REID) aims to match pedestrian images between the daytime visible and nighttime infrared camera views. The large cross-modality discrepancies have become the bottleneck which limits the performance of VI-REID. Existing methods mainly focus on capturing c…

Cited by 178PDFScholar
2021

Towards Defending against Adversarial Examples via Attack-Invariant Features

ICML 2021spotlight

Deep neural networks (DNNs) are vulnerable to adversarial noise. Their adversarial robustness can be improved by exploiting adversarial examples. However, given the continuously evolving attacks, models trained on seen types of adversarial examples generally cannot generalize well to unseen types of…

2021

Training Binary Neural Network without Batch Normalization for Image Super-Resolution

AAAI 2021technical

Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces…

Cited by 46SourcePDFScholar
2020

Binarized Neural Network for Single Image Super Resolution

ECCV 2020poster

Lighter model and faster inference are the focus of current single image super-resolution (SISR) research. However, existing methods are still hard to be applied in real-world applications due to the requirement of its heavy computation. Model quantization is an effective way to significantly reduce…

Cited by 90SourcePDFScholar
2020

Part-dependent Label Noise: Towards Instance-dependent Label Noise

NeurIPS 2020spotlight

Learning with the \textit{instance-dependent} label noise is challenging, because it is hard to model such real-world noise. Note that there are psychological and physiological evidences showing that we humans perceive instances by decomposing them into parts. Annotators are therefore more likely to…

2019

Are Anchor Points Really Indispensable in Label-Noise Learning?

NeurIPS 2019poster

In label-noise learning, the \textit{noise transition matrix}, denoting the probabilities that clean labels flip into noisy labels, plays a central role in building \textit{statistically consistent classifiers}. Existing theories have shown that the transition matrix can be learned by exploiting \te…