← Search

Xinbo Gao

99 accepted papers

2026

Dynamic-Static Collaboration for Unsupervised Domain Adaptive Video-Based Visible-Infrared Person Re-Identification

AAAI 2026technical

Video-based visible-infrared person re-identification (VVI-ReID) aims to match pedestrian sequences across modalities for all-day surveillance. While supervised methods have shown progress, their dependence on large-scale cross-modal annotations limits scalability. We investigate the task of unsuper

Cited by 2SourcePDFScholar
2026

Hyperbolic Hierarchical Alignment for Video-Based Visible-Infrared Person Re-Identification

ICML 2026poster

Video-based visible-infrared person re-identification (VVI-ReID) aims to learn robust video-level representations under modality discrepancy. However, existing methods typically rely on Euclidean geometry, which is suboptimal for modeling the complex temporal dynamics within visible and infrared tra…

Cited by 0SourceScholar
2026

Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

CVPR 2026

Incremental Object Detection (IOD) enables AI systems to continuously acquire new object classes while preserving knowledge of previously learned ones, an ability essential for deployment in dynamic, real-world environments. Existing IOD methods typically rely on knowledge distillation to mitigate c

Cited by 0SourceScholar
2026

Interference-Isolated Elastic Weight Consolidation and Knowledge Calibration for Incremental Object Detection

ICLR 2026poster

Incremental Object Detection (IOD) enables AI systems to continuously learn new object classes over time while retaining knowledge of previously learned categories. This capability is essential for adapting to dynamic environments without forgetting prior information. Although existing IOD methods h…

Cited by 0SourceScholar
2026

Learning to Watch: Active Video Anomaly Understanding via Interleaved Policy Optimization

ICML 2026poster

Video anomaly understanding (VAU) relies on sparse, context-dependent cues. However, existing passive paradigms suffer from observational aliasing, where static sampling fails to disambiguate semantically distinct events. To overcome this, we propose $Anom\text{-}\pi$, a closed-loop framework that r…

Cited by 0SourceScholar
2026

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

ICML 2026poster

Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale annotations or expert-designed priors, limiting their ability to acquire anomaly knowledge with as little human intervention as possible. To address t…

Cited by 0SourceScholar
2026

Mixture of Ranks with Degradation-Aware Routing for One-Step Real-World Image Super-Resolution

AAAI 2026technical

The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image super-resolution (Real-ISR), existing approaches mainly rely on fine-t

Cited by 0SourcePDFScholar
2026

PSMix: Robust Point Cloud Recognition through Spectral Domain Mixing

ICML 2026poster

While data augmentation is essential for robust point cloud recognition, conventional spatial mixup strategies often compromise geometric integrity by generating physically unrealistic samples. To overcome this limitation, we propose PSMix, which shifts the mixing paradigm to the spectral domain via…

Cited by 0SourceScholar
2026

R2-LIO: Real-Time and Robust LiDAR-Inertial Odometry in Dynamic Environments

ICRA 2026poster

LiDAR-Inertial Odometry (LIO) is crucial for robot navigation and autonomous driving. Most existing methods rely on the assumption of a static environment, indiscriminately using all LiDAR measurements for localization. However, LiDAR data acquired in urban scenes often contain dynamic objects such …

Cited by 0Scholar
2026

Robust Cross-Modal Retrieval via Generative Semantic Refinement and Exclusion-Guided Adaptation

ICML 2026poster

Vision-Language Pre-trained (VLP) models are vulnerable to real-world query noise. Current cross-modal Test-Time Adaptation (TTA) methods often rely on high-confidence predictions, which induces confirmation bias and neglects the informative signals in ambiguous Low-Confidence Queries. To address th…

Cited by 0SourceScholar
2026

SPR$^2$Q: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution

ICLR 2026poster

Low-bit quantization has achieved significant progress in image super-resolution. However, existing quantization methods show evident limitations in handling the heterogeneity of different components. Particularly under extreme low-bit compression, the issue of information loss becomes especially pr…

Cited by 0SourceScholar
2026

Shrinking the Teacher: An Adaptive Teaching Paradigm for Asymmetric EEG-Vision Alignment

AAAI 2026technical

Decoding visual features from EEG signals is a central challenge in neuroscience, with cross-modal alignment as the dominant approach. We argue that the relationship between visual and brain modalities is fundamentally asymmetric, characterized by two critical gaps: a Fidelity Gap (stemming from EEG

Cited by 0SourcePDFScholar
2026

StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning

ICLR 2026poster

Video Class-Incremental Learning (VCIL) seeks to develop models that continuously learn new action categories over time without forgetting previously acquired knowledge. Unlike traditional Class-Incremental Learning (CIL), VCIL introduces the added complexity of spatiotemporal structures, making it…

Cited by 0SourceScholar
2026

Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection

ICML 2026poster

Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing methods often adopt a class incremental learning perspective, separating feature spaces to sharpen decision boundaries. However, this paradigm conflicts with the inhe…

Cited by 0SourceScholar
2026

Task-Driven Subspace Decomposition for Knowledge Sharing and Isolation in LoRA-based Continual Learning

ICML 2026poster

Continual Learning (CL) requires models to sequentially adapt to new tasks without forgetting old knowledge. Recently, Low-Rank Adaptation (LoRA), a representative Parameter-Efficient Fine-Tuning (PEFT) method, has gained increasing attention in CL. Several LoRA-based CL methods reduce interference …

Cited by 0SourceScholar
2026

Towards Trustworthy Video Anomaly Understanding: A Class-Guided Chain-of-Evaluation Metric and An Anomaly-focused Meta-Benchmark

ICML 2026poster

The trustworthiness of evaluation is critical to reliable model comparison and deployment in Video Anomaly Understanding (VAU). However, existing metrics are sensitive to expression styles and normal content, and this field lacks a diagnostic benchmark to validate metric validity and robustness. To …

Cited by 0SourceScholar
2026

TransLiDAR: A Dataset and Benchmark for Cross-Sensor Point Cloud Translation

RA-L 2026

Autonomous vehicles are typically equipped with one primary and several auxiliary LiDAR sensors to generate point clouds of the environment. However, differences in structural design, resolution, and scanning mechanisms among LiDAR types lead to significant modality gaps, which hinder cross-sensor a

Cited by 0SourceScholar
2026

When Lines Meet Textures: Spatial-Frequency Aligned Diffusion Features for Cross-Sparsity Correspondence

CVPR 2026

Establishing accurate correspondence between sparse line representations and rich textured imagery remains a formidable challenge. While diffusion features excel in semantic correspondence, they struggle to bridge the fundamental gap between abstract sketches and texture-rich photographs. We identif

Cited by 0SourcecodeScholar
2025

A Multi-annotated and Multi-modal Dataset for Wide-angle Video Quality Assessment

ICASSP 2025accepted

Wide-angle video is favored for its wide viewing angle and ability to capture a large area of scenery, making it an ideal choice for sports and adventure recording. However, wide-angle video is prone to deformation, exposure and other distortions, resulting in poor video quality and affecting the pe…

Cited by 0SourceScholar
2025

A Two-Stage AIGC Image Quality Assessment with T2I Correspondence and Visual Perception

ICASSP 2025accepted

Image quality assessment (IQA) of artificial intelligence-generated content (AIGC) has recently attracted significant research attention. Unlike general-purpose IQA, which primarily focuses on evaluating image content, AIGCIQA often requires addressing both the Text-to-Image (T2I) correspondence and…

Cited by 0SourceScholar
2025

A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding

NeurIPS 2025poster

While unmanned aerial vehicles (UAVs) offer wide-area, high-altitude coverage for anomaly detection, they face challenges such as dynamic viewpoints, scale variations, and complex scenes. Existing datasets and methods, mainly designed for fixed ground-level views, struggle to adapt to these conditio…

Cited by 0SourcecodeScholar
2025

AGIAA-2K: A Fine-grained Dataset for Aesthetic and Alignment Evaluation of AI-Generated Images

ICASSP 2025accepted

With the advancement of AI-generated content technologies, AI-generated images (AGIs) have become increasingly influential in artistic creation and visual communication. However, the aesthetic quality of AGIs varies significantly due to technical limitations and the influence of user input, undersco…

Cited by 0SourceScholar
2025

Adversarial Domain Prompt Tuning and Generation for Single Domain Generalization

CVPR 2025poster

Single domain generalization (SDG) aims to learn a robust model, which could perform well on many unseen domains while there is only one single domain available for training. One of the promising directions for achieving single-domain generalization is to generate out-of-domain (OOD) training data t…

Cited by 0SourcePDFScholar
2025

Bidirectional Reference Image Quality Assessment via Content-Quality Correlation Modeling

ICASSP 2025accepted

The emphasis on no-reference image quality assessment has often overshadowed the significance of Full-Reference Image Quality Assessment (FR-IQA), which generally better reflects human contrastive perception mechanism. However, FRIQA presents challenges in obtaining content-aligned reference images.…

Cited by 0SourceScholar
2025

CognitionCapturer: Decoding Visual Stimuli from Human EEG Signal with Multimodal Information

AAAI 2025technical

Electroencephalogram (EEG) signals have attracted significant attention from researchers due to their non-invasive nature and high temporal sensitivity in decoding visual stimuli. However, most recent studies have focused solely on the relationship between EEG and image data pairs, neglecting the va…

2025

Diff-MoE: Diffusion Transformer with Time-Aware and Space-Adaptive Experts

ICML 2025poster

Diffusion models have transformed generative modeling but suffer from scalability limitations due to computational overhead and inflexible architectures that process all generative stages and tokens uniformly. In this work, we introduce Diff-MoE, a novel framework that combines Diffusion Transformer…

Cited by 0SourcePDFScholar
2025

Effective Diffusion Transformer Architecture for Image Super-Resolution

AAAI 2025technical

Recent advances indicate that diffusion model holds great promise in image super-resolution. While latest methods are primarily based on latent diffusion models with convolutional neural networks, there are few attempts to explore transformers, which have demonstrated remarkable performance in image…

2025

Information Entropy-assisted Hierarchical Framework for Unknown Environments Exploration

IROS 2025

Autonomous exploration of unknown environments is a critical task in robotic search and rescue operations. Recently, hierarchical planning frameworks have gained significant attention for their potential to enhance exploration efficiency. However, most existing approaches struggle with efficient exp

Cited by 0SourceScholar
2025

Infrared and Visible Image Fusion with Hierarchical Human Perception

ICASSP 2025accepted

Image fusion combines images from multiple domains into one image, containing complementary information from source domains. Existing methods take pixel intensity, texture and high-level vision task information as the standards to determine preservation of information, lacking enhancement for human…

Cited by 0SourceScholar
2025

MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection

NeurIPS 2025poster

Advances in image generation technologies have raised growing concerns about their potential misuse, particularly in producing misinformation and deepfakes. This creates an urgent demand for effective methods to detect AI-generated images (AIGIs). While progress has been made, achieving reliable per…

Cited by 0SourceScholar
2025

Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement

AAAI 2025technical

Adversarial fine-tuning methods enhance adversarial robustness via fine-tuning the pre-trained model in an adversarial training manner. However, we identify that some specific latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap b…

Cited by 0SourcePDFScholar
2025

Motion Artifact Removal in Pixel-Frequency Domain via Alternate Masks and Diffusion Model

AAAI 2025technical

Motion artifacts present in magnetic resonance imaging (MRI) can seriously interfere with clinical diagnosis. Removing motion artifacts is a straightforward solution and has been extensively studied. However, paired data are still heavily relied on in recent works and the perturbations in k-space (f…

2025

SAIST: Segment Any Infrared Small Target Model Guided by Contrastive Language-Image Pretraining

CVPR 2025poster

Infrared Small Target Detection (IRSTD) aims to identify low signal-to-noise ratio small targets in infrared images with complex backgrounds, which is crucial for various applications. However, existing IRSTD methods typically rely solely on image modalities for processing, which fail to fully captu…

Cited by 0SourcePDFScholar
2025

Structure-Aware Handwritten Text Recognition via Graph-Enhanced Cross-Modal Mutual Learning

IJCAI 2025

Existing handwriting recognition methods only focus on learning visual patterns by modeling low-level relationships of adjacent pixels, while overlooking the intrinsic geometric structures of characters. In this paper, we propose a novel graph-enhanced cross-modal mutual learning network GCM to full

Cited by 0SourcePDFScholar
2025

Surrogate Prompt Learning: Towards Efficient and Diverse Prompt Learning for Vision-Language Models

ICML 2025poster

Prompt learning is a cutting-edge parameter-efficient fine-tuning technique for pre-trained vision-language models (VLMs). Instead of learning a single text prompt, recent works have revealed that learning diverse text prompts can effectively boost the performances on downstream tasks, as the divers…

Cited by 0SourcePDFScholar
2025

Thinking Racial Bias in Fair Forgery Detection: Models, Datasets and Evaluations

AAAI 2025technical

Due to the successful development of deep image generation technology, forgery detection plays a more important role in social and economic security. Racial bias has not been explored thoroughly in the deep forgery detection field. In the paper, we first contribute a dedicated dataset called the Fai…

2024

Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion Model

AAAI 2024technical

Adversarial attacks involve adding perturbations to the source image to cause misclassification by the target model, which demonstrates the potential of attacking face recognition models. Existing adversarial face image generation methods still can’t achieve satisfactory performance because of low t…

2024

Beyond Euclidean: Dual-Space Representation Learning for Weakly Supervised Video Violence Detection

NeurIPS 2024poster

While numerous Video Violence Detection (VVD) methods have focused on representation learning in Euclidean space, they struggle to learn sufficiently discriminative features, leading to weaknesses in recognizing normal events that are visually similar to violent events (i.e., ambiguous violence). In…

Cited by 3SourcePDFScholar
2024

Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors

IJCAI 2024poster

The remarkable prowess of diffusion models in image generation has spurred efforts to extend their application beyond generative tasks. However, a persistent challenge exists in lacking a unified approach to apply diffusion models to visual perception tasks with diverse semantic granularity requirem…

Cited by 3SourcePDFScholar
2024

DAMSDet: Dynamic Adaptive Multispectral Detection Transformer with Competitive Query Selection and Adaptive Feature Fusion

ECCV 2024poster

"Infrared-visible object detection aims to achieve robust even full-day object detection by fusing the complementary information of infrared and visible images. However, highly dynamically variable complementary characteristics and commonly existing modality misalignment make the fusion of complemen…

2024

Diffusion-based Layer-wise Semantic Reconstruction for Unsupervised Out-of-Distribution Detection

NeurIPS 2024poster

Unsupervised out-of-distribution (OOD) detection aims to identify out-of-domain data by learning only from unlabeled In-Distribution (ID) training samples, which is crucial for developing a safe real-world machine learning system. Current reconstruction-based method provides a good alternative appro…

2024

Disentangled Prompt Representation for Domain Generalization

CVPR 2024poster

Domain Generalization (DG) aims to develop a versatile model capable of performing well on unseen target domains. Recent advancements in pre-trained Visual Foundation Models (VFMs) such as CLIP show significant potential in enhancing the generalization abilities of deep models. Although there is a g…

Cited by 9SourcePDFScholar
2024

Facial Aesthetic Enhancement Network for Asian Faces Based on Differential Facial Aesthetic Activations

ICASSP 2024accepted

In this paper, we addressed facial aesthetic enhancement (FAE). Although existing methods have made great progress, the beautified images generated by them are highly prone to poor beautification, which limits their application to real-world scenes. To tackle this problem, we proposed a new method c…

Cited by 0SourceScholar
2024

Feature-Level Adversarial Attacks and Ranking Disruption for Visible-Infrared Person Re-identification

NeurIPS 2024poster

Visible-infrared person re-identification (VIReID) is widely used in fields such as video surveillance and intelligent transportation, imposing higher demands on model security. In practice, the adversarial attacks based on VIReID aim to disrupt output ranking and quantify the security risks of mode…

Cited by 1SourcePDFScholar
2024

IRPruneDet: Efficient Infrared Small Target Detection via Wavelet Structure-Regularized Soft Channel Pruning

AAAI 2024technical

Infrared Small Target Detection (IRSTD) refers to detecting faint targets in infrared images, which has achieved notable progress with the advent of deep learning. However, the drive for improved detection accuracy has led to larger, intricate models with redundant parameters, causing storage and co…

2024

IRSAM: Advancing Segment Anything Model for Infrared Small Target Detection

ECCV 2024poster

"The recent Segment Anything Model (SAM) is a significant advancement in natural image segmentation, exhibiting potent zero-shot performance suitable for various downstream image segmentation tasks. However, directly utilizing the pretrained SAM for Infrared Small Target Detection (IRSTD) task falls…

2024

MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval

ICASSP 2024accepted

Text-to-image person retrieval aims to recognize target pedestrians based on specified text. Existing methods mainly obtain image and text features separately through distinct feature extractors, subsequently embedding them into a unified feature space and calculating their similarity. Despite great…

Cited by 0SourceScholar
2024

Multi-Granularity Graph-Convolution-Based Method for Weakly Supervised Person Search

IJCAI 2024poster

One-step Weakly Supervised Person Search (WSPS) jointly performs pedestrian detection and person Re-IDentification (ReID) only with bounding box annotations, which makes the traditional person ReID problem more suitable and efficient for real-world applications. However, this task is very challengin…

Cited by 0SourcePDFScholar
2024

Multi-Scene Generalized Trajectory Global Graph Solver with Composite Nodes for Multiple Object Tracking

AAAI 2024technical

The global multi-object tracking (MOT) system can consider interaction, occlusion, and other ``visual blur'' scenarios to ensure effective object tracking in long videos. Among them, graph-based tracking-by-detection paradigms achieve surprising performance. However, their fully-connected nature pos…

Cited by 4SourcePDFScholar
2024

On the Analysis of GAN-based Image-to-Image Translation with Gaussian Noise Injection

ICLR 2024poster

Image-to-image (I2I) translation is vital in computer vision tasks like style transfer and domain adaptation. While recent advances in GAN have enabled high-quality sample generation, real-world challenges such as noise and distortion remain significant obstacles. Although Gaussian noise injection d…

Cited by 2SourcePDFScholar
2024

Point Deformable Network with Enhanced Normal Embedding for Point Cloud Analysis

AAAI 2024technical

Recently MLP-based methods have shown strong performance in point cloud analysis. Simple MLP architectures are able to learn geometric features in local point groups yet fail to model long-range dependencies directly. In this paper, we propose Point Deformable Network (PDNet), a concise MLP-based ne…

Cited by 3SourcePDFScholar
2024

Pro2SAM: Mask Prompt to SAM with Grid Points for Weakly Supervised Object Localization

ECCV 2024poster

"Weakly Supervised Object Localization (WSOL), which aims to localize objects by only using image-level labels, has attracted much attention because of its low annotation cost in real applications. Current studies focus on the Class Activation Map (CAM) of CNN and the self-attention map of transform…

Cited by 1SourcePDFScholar
2024

SUMix: Mixup with Semantic and Uncertain Information

ECCV 2024poster

"Mixup data augmentation approaches have been applied for various tasks of deep learning to improve the generalization ability of deep neural networks. Some existing approaches CutMix, SaliencyMix, etc. randomly replace a patch in one image with patches from another to generate the mixed image. Simi…

2024

Structure-Aware in-Air Handwritten Text Recognition with Graph-Guided Cross-Modality Translator

ICASSP 2024accepted

In-air handwriting as a new human-computer interaction way plays an important role in many virtual/mixed-reality applications. Existing methods for in-air handwritten text recognition (IAHTR) typically directly process handwriting trajectories with deep neural networks. However, those methods all si…

Cited by 0SourceScholar
2024

Task-aware Orthogonal Sparse Network for Exploring Shared Knowledge in Continual Learning

ICML 2024poster

Continual learning (CL) aims to learn from sequentially arriving tasks without catastrophic forgetting (CF). By partitioning the network into two parts based on the Lottery Ticket Hypothesis---one for holding the knowledge of the old tasks while the other for learning the knowledge of the new task--…

Cited by 7SourcePDFScholar
2023

Boosting Weakly-Supervised Temporal Action Localization With Text Information

CVPR 2023poster

Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost WTAL from two aspects, i.e., (a) the discriminative objecti…

2023

Cross-Modality Person Re-identification with Memory-Based Contrastive Embedding

AAAI 2023technical

Visible-infrared person re-identification (VI-ReID) aims to retrieve the person images of the same identity from the RGB to infrared image space, which is very important for real-world surveillance system. In practice, VI-ReID is more challenging due to the heterogeneous modality discrepancy, which…

Cited by 14SourcePDFScholar
2023

DLBD: A Self-Supervised Direct-Learned Binary Descriptor

CVPR 2023poster

For learning-based binary descriptors, the binarization process has not been well addressed. The reason is that the binarization blocks gradient back-propagation. Existing learning-based binary descriptors learn real-valued output, and then it is converted to binary descriptors by their proposed bin…

2023

DecomFormer: Decompose Self-Attention Via Fourier Transform for VHR Aerial Image Scene Classification

ICASSP 2023accepted

Very high-resolution (VHR) aerial image scene classification is an essential task for aerial image understanding. Although transformer-based models have demonstrated strong ability in natural image classification, transformer-based methods on VHR aerial image tasks are still lack of concern because…

Cited by 0SourceScholar
2023

ESSAformer: Efficient Transformer for Hyperspectral Image Super-resolution

ICCV 2023poster

Single hyperspectral image super-resolution (single-HSI-SR) aims to restore a high-resolution hyperspectral image from a low-resolution observation. However, the prevailing CNN-based approaches have shown limitations in building long-range dependencies and capturing interaction information between s…

Cited by 83PDFcodeScholar
2023

Eliminating Adversarial Noise via Information Discard and Robust Representation Restoration

ICML 2023poster

Deep neural networks (DNNs) are vulnerable to adversarial noise. Denoising model-based defense is a major protection strategy. However, denoising models may fail and induce negative effects in fully white-box scenarios. In this work, we start from the latent inherent properties of adversarial sample…

Cited by 8SourcePDFScholar
2023

FCIR: Rethink Aerial Image Super Resolution with Fourier Analysis

ICASSP 2023accepted

Recent years, deep-learning-based methods achieve remarkable improvements on the super-resolution (SR) task. However, recovering high-quality (HQ) texture from the low-quality (LQ) aerial image is still challenging due to the limited contextual modeling ability of current deep-learning methods as we…

Cited by 0SourceScholar
2023

Hiding Visual Information via Obfuscating Adversarial Perturbations

ICCV 2023poster

Growing leakage and misuse of visual information raise security and privacy concerns, which promotes the development of information protection. Existing adversarial perturbations-based methods mainly focus on the de-identification against deep learning models. However, the inherent visual informatio…

Cited by 12PDFcodeScholar
2023

Hierarchical Point-based Active Learning for Semi-supervised Point Cloud Semantic Segmentation

ICCV 2023poster

Impressive performance on point cloud semantic segmentation has been achieved by fully-supervised methods with large amounts of labelled data. As it is labour-intensive to acquire large-scale point cloud data with point-wise labels, many attempts have been made to explore learning 3D point cloud seg…

Cited by 19PDFcodeScholar
2023

Hierarchical Supervision and Shuffle Data Augmentation for 3D Semi-Supervised Object Detection

CVPR 2023poster

State-of-the-art 3D object detectors are usually trained on large-scale datasets with high-quality 3D annotations. However, such 3D annotations are often expensive and time-consuming, which may not be practical for real applications. A natural remedy is to adopt semi-supervised learning (SSL) by lev…

2023

MCF: Mutual Correction Framework for Semi-Supervised Medical Image Segmentation

CVPR 2023poster

Semi-supervised learning is a promising method for medical image segmentation under limited annotation. However, the model cognitive bias impairs the segmentation performance, especially for edge regions. Furthermore, current mainstream semi-supervised medical image segmentation (SSMIS) methods lack…

2023

Phase-aware Adversarial Defense for Improving Adversarial Robustness

ICML 2023poster

Deep neural networks have been found to be vulnerable to adversarial noise. Recent works show that exploring the impact of adversarial noise on intrinsic components of data can help improve adversarial robustness. However, the pattern closely related to human perception has not been deeply studied.…

Cited by 8SourcePDFScholar
2023

Self-Supervised Image Local Forgery Detection by JPEG Compression Trace

AAAI 2023technical

For image local forgery detection, the existing methods require a large amount of labeled data for training, and most of them cannot detect multiple types of forgery simultaneously. In this paper, we firstly analyzed the JPEG compression traces which are mainly caused by different JPEG compression c…

Cited by 7SourcePDFScholar
2022

Boundary-Aware Bias Loss for Transformer-Based Aerial Image Segmentation Model

ICASSP 2022accepted

Inspired by the tremendous success of the transformer-based model in natural language processing (NLP), many efforts introduce the transformer-based model into the image processing tasks. However, naive transformer models have to down-sample the image resolution to satisfy computational restrictions…

Cited by 0SourceScholar
2022

Class-Dependent Label-Noise Learning with Cycle-Consistency Regularization

NeurIPS 2022accept

In label-noise learning, estimating the transition matrix plays an important role in building statistically consistent classifier. Current state-of-the-art consistent estimator for the transition matrix has been developed under the newly proposed sufficiently scattered assumption, through incorporat…

Cited by 40SourcePDFScholar
2022

Improving Adversarial Robustness via Mutual Information Estimation

ICML 2022spotlight

Deep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples f…

2022

Instance-Dependent Label-Noise Learning With Manifold-Regularized Transition Matrix Estimation

CVPR 2022poster

In label-noise learning, estimating the transition matrix has attracted more and more attention as the matrix plays an important role in building statistically consistent classifiers. However, it is very challenging to estimate the transition matrix T(x), where T(x) denotes the instance, because it…

Cited by 93PDFScholar
2022

Robust Single Image Dehazing Based on Consistent and Contrast-Assisted Reconstruction

IJCAI 2022poster

Single image dehazing as a fundamental low-level vision task, is essential for the development of robust intelligent surveillance system. In this paper, we make an early effort to consider dehazing robustness under variational haze density, which is a realistic while under-studied problem in the res…

Cited by 7SourcePDFScholar
2022

SS3D: Sparsely-Supervised 3D Object Detection From Point Cloud

CVPR 2022poster

Conventional deep learning based methods for 3D object detection require a large amount of 3D bounding box annotations for training, which is expensive to obtain in practice. Sparsely annotated object detection, which can largely reduce the annotations, is very challenging since the missingannotated…

Cited by 28PDFcodeScholar
2022

Towards Semi-Supervised Deep Facial Expression Recognition With an Adaptive Confidence Margin

CVPR 2022poster

Only parts of unlabeled data are selected to train models for most semi-supervised learning methods, whose confidence scores are usually higher than the pre-defined threshold (i.e., the confidence margin). We argue that the recognition performance should be further improved by making full use of all…

Cited by 116PDFcodeScholar
2021

A Circular-Structured Representation for Visual Emotion Distribution Learning

CVPR 2021poster

Visual Emotion Analysis (VEA) has attracted increasing attention recently with the prevalence of sharing images on social networks. Since human emotions are ambiguous and subjective, it is more reasonable to address VEA in a label distribution learning (LDL) paradigm rather than a single-label class…

Cited by 39PDFScholar
2021

A Sketch-Transformer Network for Face Photo-Sketch Synthesis

IJCAI 2021poster

We present a face photo-sketch synthesis model, which converts a face photo into an artistic face sketch or recover a photo-realistic facial image from a sketch portrait. Recent progress has been made by convolutional neural networks (CNNs) and generative adversarial networks (GANs), so that promisi…

Cited by 32SourcePDFScholar
2021

Drafting and Revision: Laplacian Pyramid Network for Fast High-Quality Artistic Style Transfer

CVPR 2021poster

Artistic style transfer aims at migrating the style from an example image to a content image. Currently, optimization-based methods have achieved great stylization quality, but expensive time cost restricts their practical applications. Meanwhile, feed-forward methods still fail to synthesize comple…

Cited by 119PDFcodeScholar
2021

Removing Adversarial Noise in Class Activation Feature Space

ICCV 2021poster

Deep neural networks (DNNs) are vulnerable to adversarial noise. Pre-processing based defenses could largely remove adversarial noise by processing inputs. However, they are typically affected by the error amplification effect, especially in the front of continuously evolving attacks. To solve this…

Cited by 36PDFcodeScholar
2021

Support-Set Based Cross-Supervision for Video Grounding

ICCV 2021poster

Current approaches for video grounding propose kinds of complex architectures to capture the video-text relations, and have achieved impressive improvements. However, it is hard to learn the complicated multi-modal relations by only architecture designing in fact. In this paper, we introduce a novel…

Cited by 53PDFScholar
2021

Syncretic Modality Collaborative Learning for Visible Infrared Person Re-Identification

ICCV 2021poster

Visible infrared person re-identification (VI-REID) aims to match pedestrian images between the daytime visible and nighttime infrared camera views. The large cross-modality discrepancies have become the bottleneck which limits the performance of VI-REID. Existing methods mainly focus on capturing c…

Cited by 178PDFScholar
2021

TSGCNet: Discriminative Geometric Feature Learning With Two-Stream Graph Convolutional Network for 3D Dental Model Segmentation

CVPR 2021poster

The ability to segment teeth precisely from digitized 3D dental models is an essential task in computer-aided orthodontic surgical planning. To date, deep learning based methods have been popularly used to handle this task. State-of-the-art methods directly concatenate the raw attributes of 3D input…

Cited by 55PDFcodeScholar
2021

Towards Defending against Adversarial Examples via Attack-Invariant Features

ICML 2021spotlight

Deep neural networks (DNNs) are vulnerable to adversarial noise. Their adversarial robustness can be improved by exploiting adversarial examples. However, given the continuously evolving attacks, models trained on seen types of adversarial examples generally cannot generalize well to unseen types of…

2021

Training Binary Neural Network without Batch Normalization for Image Super-Resolution

AAAI 2021technical

Recently, binary neural network (BNN) based super-resolution (SR) methods have enjoyed initial success in the SR field. However, there is a noticeable performance gap between the binarized model and the full-precision one. Furthermore, the batch normalization (BN) in binary SR networks introduces…

Cited by 46SourcePDFScholar
2020

Binarized Neural Network for Single Image Super Resolution

ECCV 2020poster

Lighter model and faster inference are the focus of current single image super-resolution (SISR) research. However, existing methods are still hard to be applied in real-world applications due to the requirement of its heavy computation. Model quantization is an effective way to significantly reduce…

Cited by 90SourcePDFScholar
2018

Blind Image Quality Assessment Based on Visuo-Spatial Series Statistics

ICASSP 2018accepted

Existing blind image quality assessment (BIQA) methods based on statistics attach limited attention to the relative position of pixels. Features in these BIQA methods are too flimsy to characterize quite a few distortions with strong locality or complexity. However, psychological studies have shown…

Cited by 0SourceScholar
2018

Fast and Accurate Single Image Super-Resolution via Information Distillation Network

CVPR 2018poster

Recently, deep convolutional neural networks (CNNs) have been demonstrated remarkable progress on single image super-resolution. However, as the depth and width of the networks increase, CNN-based super-resolution methods have been faced with the challenges of computational complexity and memory con…

2018

Self-Supervised Adversarial Hashing Networks for Cross-Modal Retrieval

CVPR 2018poster

Thanks to the success of deep learning, cross-modal retrieval has made significant progress recently. However, there still remains a crucial bottleneck: how to bridge the modality gap to further enhance the retrieval accuracy. In this paper, we propose a self-supervised adversarial hashing (SSAH) ap…

2015

Coupled fisher discrimination dictionary learning for single image super-resolution

ICASSP 2015accepted

Image Super-resolution (SR) reconstruction techniques based on sparse representation have attracted ever-increasing attentions in recent years, where the choice of over-complete dictionary is of prime important for reconstruction quality. However, most of the image SR methods based on sparse represe…

Cited by 0SourceScholar