← Search

Xian-Sheng Hua

72 accepted papers

2026

CLINIC: Towards High-quality Graph Out-Of-Distribution Detection

ICML 2026poster

This paper studies the problem of graph out-of-distribution (OOD) detection, which aims to identify anomaly graphs out of a graph dataset. Prior efforts usually focus on the utilization of topological structures with unsupervised graph learning to foster typical pattern recognition, which overlooks …

Cited by 0SourceScholar
2026

Intrinsic Gradient Suppression for Label-Noise Prompt Tuning in Vision–Language Models

ICML 2026poster

Contrastive vision-language models like CLIP exhibit remarkable zero-shot generalization. However, prompt tuning remains highly sensitive to label noise, as mislabeled samples generate disproportionately large gradients that can overwhelm pre-trained priors. We argue that because CLIP already provid…

Cited by 0SourceScholar
2025

DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal Retrieval

AAAI 2025technical

With the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, wh…

Cited by 0SourcePDFScholar
2025

LEAF: Large Language Diffusion Model for Time Series Forecasting

EMNLP 2025

This paper studies the problem of time series forecasting, which aims to generate future predictions given historical trajectories. Recent researchers have applied large language models (LLMs) into time series forecasting, which usually align the time series space with textual space and output futur

Cited by 0SourcePDFScholar
2025

POPoS: Improving Efficient and Robust Facial Landmark Detection with Parallel Optimal Position Search

AAAI 2025technical

Achieving a balance between accuracy and efficiency is a critical challenge in facial landmark detection (FLD). This paper introduces Parallel Optimal Position Search (POPoS), a high-precision encoding-decoding framework designed to address the limitations of traditional FLD methods. POPoS employs t…

2025

SDBench: A Survey-based Domain-specific LLM Benchmarking and Optimization Framework

ACL 2025long

The rapid advancement of large language models (LLMs) in recent years has made it feasible to establish domain-specific LLMs for specialized fields. However, in practical development, acquiring domain-specific knowledge often requires a significant amount of professional expert manpower. Moreover, e…

Cited by 0SourcePDFScholar
2025

scRAG: Hybrid Retrieval-Augmented Generation for LLM-based Cross-Tissue Single-Cell Annotation

ACL 2025finding

In recent years, large language models (LLMs) such as GPT-4 have demonstrated impressive potential in a wide range of fields, including biology, genomics and healthcare. Numerous studies have attempted to apply pre-trained LLMs to single-cell data analysis within one tissue. However, when it comes t…

2024

BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition

CVPR 2024poster

Graph Convolutional Networks (GCNs) have long set the state-of-the-art in skeleton-based action recognition leveraging their ability to unravel the complex dynamics of human joint topology through the graph's adjacency matrix. However an inherent flaw has come to light in these cutting-edge models:…

2024

Grid-Attention: Enhancing Computational Efficiency of Large Vision Models without Fine-Tuning

ECCV 2024poster

"Recently, transformer-based large vision models, , the Segment Anything Model (SAM) and Stable Diffusion (SD), have achieved remarkable success in the computer vision field. However, the quartic complexity within the transformer’s Multi-Head Attention (MHA) leads to substantial computational costs…

2024

PURE: Prompt Evolution with Graph ODE for Out-of-distribution Fluid Dynamics Modeling

NeurIPS 2024poster

This work studies the problem of out-of-distribution fluid dynamics modeling. Previous works usually design effective neural operators to learn from mesh-based data structures. However, in real-world applications, they would suffer from distribution shifts from the variance of system parameters and…

Cited by 5SourcePDFScholar
2024

Prometheus: Out-of-distribution Fluid Dynamics Modeling with Disentangled Graph ODE

ICML 2024poster

Fluid dynamics modeling has received extensive attention in the machine learning community. Although numerous graph neural network (GNN) approaches have been proposed for this problem, the problem of out-of-distribution (OOD) generalization remains underexplored. In this work, we propose a new large…

Cited by 10SourcePDFScholar
2024

Semi-supervised Knowledge Transfer Across Multi-omic Single-cell Data

NeurIPS 2024poster

Knowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data.…

Cited by 0SourcePDFScholar
2023

CoCo: A Coupled Contrastive Framework for Unsupervised Domain Adaptive Graph Classification

ICML 2023poster

Although graph neural networks (GNNs) have achieved impressive achievements in graph classification, they often need abundant task-specific labels, which could be extensively costly to acquire. A credible solution is to explore additional labeled graphs to enhance unsupervised learning on the target…

Cited by 33SourcePDFScholar
2023

IDEA: An Invariant Perspective for Efficient Domain Adaptive Image Retrieval

NeurIPS 2023poster

In this paper, we investigate the problem of unsupervised domain adaptive hashing, which leverage knowledge from a label-rich source domain to expedite learning to hash on a label-scarce target domain. Although numerous existing approaches attempt to incorporate transfer learning techniques into dee…

Cited by 6SourcePDFScholar
2023

Invariant Training 2D-3D Joint Hard Samples for Few-Shot Point Cloud Recognition

ICCV 2023poster

We tackle the data scarcity challenge in few-shot point cloud recognition of 3D objects by using a joint prediction from a conventional 3D model and a well-pretrained 2D model. Surprisingly, such an ensemble, though seems trivial, has hardly been shown effective in recent 2D-3D models. We find out t…

Cited by 13PDFcodeScholar
2023

Prototypical Mixing and Retrieval-Based Refinement for Label Noise-Resistant Image Retrieval

ICCV 2023poster

Label noise is pervasive in real-world applications, which influences the optimization of neural network models. This paper investigates a realistic but understudied problem of image retrieval under label noise, which could lead to severe overfitting or memorization of noisy samples during optimizat…

Cited by 5PDFcodeScholar
2022

Active Boundary Loss for Semantic Segmentation

AAAI 2022technical

This paper proposes a novel active boundary loss for semantic segmentation. It can progressively encourage the alignment between predicted boundaries and ground-truth boundaries during end-to-end training, which is not explicitly enforced in commonly used cross-entropy loss. Based on the predicted b…

2022

Balanced and Hierarchical Relation Learning for One-Shot Object Detection

CVPR 2022poster

Instance-level feature matching is significantly important to the success of modern one-shot object detectors. Recently, the methods based on the metric-learning paradigm have achieved an impressive process. Most of these works only measure the relations between query and target objects on a single…

Cited by 31PDFcodeScholar
2022

Box-Supervised Instance Segmentation with Level Set Evolution

ECCV 2022poster

"In contrast to the fully supervised methods using pixel-wise mask labels, box-supervised instance segmentation takes advantage of the simple box annotations, which has recently attracted a lot of research attentions. In this paper, we propose a novel single-shot box-supervised instance segmentation…

2022

Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-of-Distribution Generalization

ECCV 2022poster

"Out-Of-Distribution generalization (OOD) is all about learning invariance against environmental changes. If the context in every class is evenly distributed, OOD would be trivial because the context can be easily removed due to an underlying principle: class is invariant to context. However, collec…

2022

Class Re-Activation Maps for Weakly-Supervised Semantic Segmentation

CVPR 2022poster

Extracting class activation maps (CAM) is arguably the most standard step of generating pseudo masks for weakly-supervised semantic segmentation (WSSS). Yet, we find that the crux of the unsatisfactory pseudo masks is the binary cross-entropy loss (BCE) widely used in CAM. Specifically, due to the s…

Cited by 214PDFcodeScholar
2022

Cloth-Changing Person Re-Identification From a Single Image With Gait Prediction and Regularization

CVPR 2022poster

Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting,…

Cited by 179PDFcodeScholar
2022

Cross-Domain Empirical Risk Minimization for Unbiased Long-Tailed Classification

AAAI 2022technical

We address the overlooked unbiasedness in existing long-tailed classification methods: we find that their overall improvement is mostly attributed to the biased preference of "tail" over "head", as the test distribution is assumed to be balanced; however, when the test is as imbalanced as the long-t…

2022

Delving into Details: Synopsis-to-Detail Networks for Video Recognition

ECCV 2022poster

"In this paper, we explore the details in video recognition with the aim to improve the accuracy. It is observed that most failure cases in recent works fall on the mis-classifications among very similar actions (such as high kick vs. side kick) that need a capturing of fine-grained discriminative d…

2022

Dense Learning Based Semi-Supervised Object Detection

CVPR 2022poster

The ultimate goal of semi-supervised object detection (SSOD) is to facilitate the utilization and deployment of detectors in actual applications with the help of a large amount of unlabeled data. Although a few works have proposed various self-training-based methods or consistency-regularization-bas…

Cited by 87PDFcodeScholar
2022

Homography Loss for Monocular 3D Object Detection

CVPR 2022poster

Monocular 3D object detection is an essential task in autonomous driving. However, most current methods consider each 3D object in the scene as an independent training sample, while ignoring their inherent geometric relations, thus inevitably resulting in a lack of leveraging spatial constraints. In…

Cited by 59PDFcodeScholar
2022

Identifying Hard Noise in Long-Tailed Sample Distribution

ECCV 2022poster

"Conventional de-noising methods rely on the assumption that the noisy samples are independent and identically distributed, so the resultant classifier, though disturbed by noise, can still easily identify the noises as outliers. However, the assumption is unrealistic in large-scale data that is ine…

2022

Meta Convolutional Neural Networks for Single Domain Generalization

CVPR 2022poster

In single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is…

Cited by 60PDFScholar
2022

On Non-Random Missing Labels in Semi-Supervised Learning

ICLR 2022poster

Semi-Supervised Learning (SSL) is fundamentally a missing label problem, in which the label Missing Not At Random (MNAR) problem is more realistic and challenging, compared to the widely-adopted yet naive Missing Completely At Random assumption where both labeled and unlabeled data share the same cl…

2022

Online Convolutional Re-Parameterization

CVPR 2022poster

Structural re-parameterization has drawn increasing attention in various computer vision tasks. It aims at improving the performance of deep models without introducing any inference-time cost. Though efficient during inference, such models rely heavily on the complicated training-time blocks to achi…

Cited by 91PDFcodeScholar
2022

Rethinking IoU-Based Optimization for Single-Stage 3D Object Detection

ECCV 2022poster

"Since Intersection-over-Union (IoU) based optimization maintains the consistency of the final IoU prediction metric and losses, it has been widely used in both regression and classification branches of single-stage 2D object detectors. Recently, several 3D object detection methods adopt IoU-based o…

2022

Spatiotemporal Self-Attention Modeling with Temporal Patch Shift for Action Recognition

ECCV 2022poster

"Transformer-based methods have recently achieved great advancement on 2D image-based vision tasks. For 3D video-based tasks such as action recognition, however, directly applying spatiotemporal transformers on video data will bring heavy computation and memory burdens due to the largely increased n…

2022

Structural and Statistical Texture Knowledge Distillation for Semantic Segmentation

CVPR 2022poster

Existing knowledge distillation works for semantic segmentation mainly focus on transfering high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the local structural pattern and global statistical property, such…

Cited by 77PDFScholar
2022

TGNN: A Joint Semi-supervised Framework for Graph-level Classification

IJCAI 2022poster

This paper studies semi-supervised graph classification, a crucial task with a wide range of applications in social network analysis and bioinformatics. Recent works typically adopt graph neural networks to learn graph-level representations for classification, failing to explicitly leverage features…

Cited by 48SourcePDFScholar
2022

Unpaired Cartoon Image Synthesis via Gated Cycle Mapping

CVPR 2022poster

In this paper, we present a general-purpose solution to cartoon image synthesis with unpaired training data. In contrast to previous works learning pre-defined cartoon styles for specified usage scenarios (portrait or scene), we aim to train a common cartoon translator which can not only simultaneou…

Cited by 21PDFScholar
2021

$\alpha$-IoU: A Family of Power Intersection over Union Losses for Bounding Box Regression

NeurIPS 2021poster

Bounding box (bbox) regression is a fundamental task in computer vision. So far, the most commonly used loss functions for bbox regression are the Intersection over Union (IoU) loss and its variants. In this paper, we generalize existing IoU-based losses to a new family of power IoU losses that have…

2021

3D Local Convolutional Neural Networks for Gait Recognition

ICCV 2021poster

The goal of gait recognition is to learn the unique spatio-temporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existi…

Cited by 134PDFcodeScholar
2021

Asynchronous Teacher Guided Bit-wise Hard Mining for Online Hashing

AAAI 2021technical

Online hashing for streaming data has attracted increasing attention recently. However, most existing algorithms focus on batch inputs and instance-balanced optimization, which is limited in the single datum input case and does not match the dynamic training in online hashing. Furthermore, constantl…

Cited by 9SourcePDFScholar
2021

Camera-Aware Proxies for Unsupervised Person Re-Identification

AAAI 2021technical

This paper tackles the purely unsupervised person re-identification (Re-ID) problem that requires no annotations. Some previous methods adopt clustering techniques to generate pseudo labels and use the produced labels to train Re-ID models progressively. These methods are relatively simple but effec…

2021

Category Dictionary Guided Unsupervised Domain Adaptation for Object Detection

AAAI 2021technical

Unsupervised domain adaption (UDA) is a promising solution to enhance the generalization ability of a model from a source domain to a target domain without manually annotating labels for target data. Recent works in cross-domain object detection mostly resort to adversarial feature adaptation to mat…

Cited by 49SourcePDFScholar
2021

Counterfactual VQA: A Cause-Effect Look at Language Bias

CVPR 2021poster

Recent VQA models may tend to rely on language bias as a shortcut and thus fail to sufficiently learn the multi-modal knowledge from both vision and language. In this paper, we investigate how to capture and mitigate language bias in VQA. Motivated by causal effects, we proposed a novel counterfactu…

Cited by 491PDFcodeScholar
2021

Counterfactual Zero-Shot and Open-Set Visual Recognition

CVPR 2021poster

We present a novel counterfactual framework for both Zero-Shot Learning (ZSL) and Open-Set Recognition (OSR), whose common challenge is generalizing to the unseen-classes by only training on the seen-classes. Our idea stems from the observation that the generated samples for unseen-classes are often…

Cited by 251PDFcodeScholar
2021

DCT-Mask: Discrete Cosine Transform Mask Representation for Instance Segmentation

CVPR 2021poster

Binary grid mask representation is broadly used in instance segmentation. A representative instantiation is Mask R-CNN which predicts masks on a 28*28 binary grid. Generally, a low-resolution grid is not sufficient to capture the details, while a high-resolution grid dramatically increases the train…

Cited by 88PDFcodeScholar
2021

Dense Interaction Learning for Video-Based Person Re-Identification

ICCV 2021poster

Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction…

Cited by 67PDFcodeScholar
2021

Distilling Causal Effect of Data in Class-Incremental Learning

CVPR 2021poster

We propose a causal framework to explain the catastrophic forgetting in Class-Incremental Learning (CIL) and then derive a novel distillation method that is orthogonal to the existing anti-forgetting techniques, such as data replay and feature/label distillation. We first 1) place CIL into the frame…

Cited by 257PDFcodeScholar
2021

Graph Contrastive Clustering

ICCV 2021poster

Recently, some contrastive learning methods have been proposed to simultaneously learn representations and clustering assignments, achieving significant improvements. However, these methods do not take the category information and clustering objective into consideration, thus the learned representat…

Cited by 170PDFcodeScholar
2021

Improving 3D Object Detection With Channel-Wise Transformer

ICCV 2021poster

Though 3D object detection from point clouds has achieved rapid progress in recent years, the lack of flexible and high-performance proposal refinement remains a great hurdle for existing state-of-the-art two-stage detectors. Previous works on refining 3D proposals have relied on human-designed comp…

Cited by 298PDFcodeScholar
2021

Interactive Self-Training With Mean Teachers for Semi-Supervised Object Detection

CVPR 2021poster

The goal of semi-supervised object detection is to learn a detection model using only a few labeled data and large amounts of unlabeled data, thereby reducing the cost of data labeling. Although a few studies have proposed various self-training-based methods or consistency regularization-based metho…

Cited by 170PDFScholar
2021

Partial Person Re-Identification With Part-Part Correspondence Learning

CVPR 2021poster

Driven by the success of deep learning, the last decade has seen rapid advances in person re-identification (re-ID). Nonetheless, most of approaches assume that the input is given with the fulfillment of expectations, while imperfect input remains rarely explored to date, which is a non-trivial prob…

Cited by 52PDFScholar
2021

Revisiting Knowledge Distillation: An Inheritance and Exploration Framework

CVPR 2021poster

Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the t…

Cited by 41PDFcodeScholar
2021

Traffic Flow Prediction with Vehicle Trajectories

AAAI 2021technical

This paper proposes a spatiotemporal deep learning framework, Trajectory-based Graph Neural Network (TrGNN), that mines the underlying causality of flows from historical vehicle trajectories and incorporates that into road traffic prediction. The vehicle trajectory transition patterns are studied to…

2021

Transporting Causal Mechanisms for Unsupervised Domain Adaptation

ICCV 2021poster

Existing Unsupervised Domain Adaptation (UDA) literature adopts the covariate shift and conditional shift assumptions, which essentially encourage models to learn common features across domains. However, due to the lack of supervision in the target domain, they suffer from the semantic loss: the fea…

Cited by 79PDFcodeScholar
2021

Video Object Segmentation With Dynamic Memory Networks and Adaptive Object Alignment

ICCV 2021poster

In this paper, we propose a novel solution for object-matching based semi-supervised video object segmentation, where the target object masks in the first frame are provided. Existing object-matching based methods focus on the matching between the raw object features of the current frame and the fir…

Cited by 35PDFcodeScholar
2020

Adversarial Mutual Information for Text Generation

ICML 2020poster

Recent advances in maximizing mutual information (MI) between the source and target have demonstrated its effectiveness in text generation. However, previous works paid little attention to modeling the backward network of MI (i.e., dependency from the target to the source), which is crucial to the t…

2020

Boosting Semantic Human Matting With Coarse Annotations

CVPR 2020oral

Semantic human matting aims to estimate the per-pixel opacity of the foreground human regions. It is quite challenging that usually requires user interactive trimaps and plenty of high quality annotated data. Annotating such kind of data is labor intensive and requires great skills beyond normal use…

Cited by 113PDFScholar
2020

CPR-GCN: Conditional Partial-Residual Graph Convolutional Network in Automated Anatomical Labeling of Coronary Arteries

CVPR 2020oral

Automated anatomical labeling plays a vital role in coronary artery disease diagnosing procedure. The main challenge in this problem is the large individual variability inherited in human anatomy. Existing methods usually rely on the position information and the prior knowledge of the topology of th…

Cited by 56PDFScholar
2020

Causal Intervention for Weakly-Supervised Semantic Segmentation

NeurIPS 2020oral

We present a causal inference framework to improve Weakly-Supervised Semantic Segmentation (WSSS). Specifically, we aim to generate better pixel-level pseudo-masks by using only image-level labels -- the most crucial step in WSSS. We attribute the cause of the ambiguous boundaries of pseudo-masks to…

2020

MaCAR: Urban Traffic Light Control via Active Multi-agent Communication and Action Rectification

IJCAI 2020poster

Urban traffic light control is an important and challenging real-world problem. By regarding intersections as agents, most of the Reinforcement Learning (RL) based methods generate actions of agents independently. They can cause action conflict and result in overflow or road resource waste in adjace…

Cited by 0SourcePDFScholar
2020

SLV: Spatial Likelihood Voting for Weakly Supervised Object Detection

CVPR 2020poster

Based on the framework of multiple instance learning (MIL), tremendous works have promoted the advances of weakly supervised object detection (WSOD). However, most MIL-based methods tend to localize instances to their discriminative parts instead of the whole content. In this paper, we propose a spa…

Cited by 95PDFScholar
2020

Structure Aware Single-Stage 3D Object Detection From Point Cloud

CVPR 2020poster

3D object detection from point cloud data plays an essential role in autonomous driving. Current single-stage detectors are efficient by progressively downscaling the 3D point clouds in a fully convolutional manner. However, the downscaled features inevitably lose spatial information and cannot make…

Cited by 720PDFcodeScholar
2019

Attribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-Identification

CVPR 2019poster

Video-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-pooling, without exploring the different frame weights caused by various viewpoi…

Cited by 184PDFScholar
2019

Dynamic Anchor Feature Selection for Single-Shot Object Detection

ICCV 2019poster

The design of anchors is critical to the performance of one-stage detectors. Recently, the anchor refinement module (ARM) has been proposed to adjust the initialization of default anchors, providing the detector a better anchor reference. However, this module brings another problem: all pixels at a…

Cited by 53PDFScholar
2018

An Adversarial Approach to Hard Triplet Generation

ECCV 2018poster

While deep neural networks have demonstrated competitive results for many visual recognition and image retrieval tasks, the major challenge lies in distinguishing similar images from different categories (i.e., hard negative examples) while clustering images with large variations from the same categ…

Cited by 121SourcePDFScholar
2018

Global Versus Localized Generative Adversarial Nets

CVPR 2018poster

In this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that {em globally} parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data poi…

Cited by 97SourcePDFScholar
2017

Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images

CVPR 2017poster

In recent years, both online retail and video hosting service have been exponentially grown. In this paper, a novel deep neural network, called AsymNet, is proposed to explore a new cross-domain task, Video2Shop, targeting for matching clothes appeared in videos to the exactly same items in online s…

Cited by 108PDFcodeScholar