← Search

Hanzi Wang

35 accepted papers

2026

Is Spurious Correlation Removal Always Learnable?

ICML 2026poster

Invariant learning can fail even when the invariant structure is statistically identifiable. We show an inherent computational barrier: under the Planted Clique hypothesis, there exist samplable linear-Gaussian multi-environment instances with a one-dimensional invariant subspace ($k=1$) that are le…

Cited by 0SourceScholar
2026

Joint Implicit and Explicit Language Learning for Pedestrian Attribute Recognition

AAAI 2026technical

Pedestrian attribute recognition (PAR) has received increasing attention due to its wide application in video surveillance and pedestrian analysis. Some text-enhanced methods tackle this task by converting attributes into language descriptions to facilitate interactive learning between attributes an

Cited by 0SourcePDFScholar
2026

Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition

CVPR 2026

Adapting Vision-Language Models (VLMs) to few-shot action recognition (FSAR) often trades accuracy for stability: task-specific gains can trigger catastrophic forgetting of domain-general knowledge and reduce inter-class margins. In few-shot episodes, each query is contrasted with only one positive

Cited by 0SourceScholar
2026

SAM2-OV: A Novel Detection-Only Tuning Paradigm for Open-Vocabulary Multi-Object Tracking

AAAI 2026technical

Open-vocabulary multi-object tracking (OV-MOT) aims to track objects with unseen categories beyond the training set. While existing methods rely on pseudo video sequences synthesized from static images, they struggle to model realistic motion patterns, resulting in limited association performance in

Cited by 0SourcePDFScholar
2025

Language Decoupling with Fine-grained Knowledge Guidance for Referring Multi-object Tracking

ICCV 2025poster

Referring Multi-Object Tracking (RMOT) aims to detect and track specific objects based on natural language expressions. Previous methods typically rely on sentence-level vision-language alignment, often failing to exploit fine-grained linguistic cues that are crucial for distinguishing objects with…

2025

Mind the Gap: Confidence Discrepancy Can Guide Federated Semi-Supervised Learning Across Pseudo-Mismatch

CVPR 2025poster

Federated Semi-Supervised Learning (FSSL) aims to leverage unlabeled data across clients with limited labeled data to train a global model with strong generalization ability. Most FSSL methods rely on consistency regularization with pseudo-labels, converting predictions from local or global models i…

2025

Unlocker: Disentangle the Deadlock of Learning between Label-noisy and Long-tailed Data

NeurIPS 2025poster

In real world, the observed label distribution of a dataset often mismatches its true distribution due to noisy labels. In this situation, noisy labels learning (NLL) methods directly integrated with long-tail learning (LTL) methods tend to fail due to a dilemma: NLL methods normally rely o…

Cited by 0SourceScholar
2025

WarpGAN: Warping-Guided 3D GAN Inversion with Style-Based Novel View Inpainting

NeurIPS 2025poster

3D GAN inversion projects a single image into the latent space of a pre-trained 3D GAN to achieve single-shot novel view synthesis, which requires visible regions with high fidelity and occluded regions with realism and multi-view consistency. However, existing methods focus on the reconstruction o…

Cited by 0SourceScholar
2025

You Are Your Own Best Teacher: Achieving Centralized-level Performance in Federated Learning under Heterogeneous and Long-tailed Data

ICCV 2025poster

Data heterogeneity, stemming from local non-IID data and global long-tailed distributions, is a major challenge in federated learning (FL), leading to significant performance gaps compared to centralized learning. Previous research found that poor representations and biased classifiers are the main…

2024

Bi-Directional Motion Attention with Contrastive Learning for few-shot Action Recognition

ICASSP 2024accepted

In recent years, many few-shot action recognition methods have achieved competitive performance by adopting metric-based techniques. However, they suffer from two limitations: (1) Spatio-temporal relationship is modeled independently, overlooking the spatio-temporal correspondence between target obj…

Cited by 0SourceScholar
2024

Dynamically Anchored Prompting for Task-Imbalanced Continual Learning

IJCAI 2024poster

Existing continual learning literature relies heavily on a strong assumption that tasks arrive with a balanced data stream, which is often unrealistic in real-world applications. In this work, we explore task-imbalanced continual learning (TICL) scenarios where the distribution of task data is non-u…

2024

Federated Learning with Extremely Noisy Clients via Negative Distillation

AAAI 2024technical

Federated learning (FL) has shown remarkable success in cooperatively training deep models, while typically struggling with noisy labels. Advanced works propose to tackle label noise by a re-weighting strategy with a strong assumption, i.e., mild label noise. However, it may be violated in many real…

2024

Proposal Distillation of Multi-Modal Feature Aggregation Network for Video Object Detection

ICASSP 2024accepted

Video object detection is a challenging task due to deteriorated object appearances. In order to bolster per-frame feature representations, one way is to aggregate features from relevant frames. However, relying exclusively on RGB modal for feature aggregation may limit the detection performance for…

Cited by 0SourceScholar
2024

Spatial-Contextual Discrepancy Information Compensation for GAN Inversion

AAAI 2024technical

Most existing GAN inversion methods either achieve accurate reconstruction but lack editability or offer strong editability at the cost of fidelity. Hence, how to balance the distortion-editability trade-off is a significant challenge for GAN inversion. To address this challenge, we introduce a nov…

2024

Spatio-Temporal Correlation Learning for Multiple Object Tracking

ICASSP 2024accepted

Multi-object tracking (MOT) has gained remarkable progress in recent years, while due to the complexity of real-world environments, there are still many challenges that remain unsolved, such as object occlusion and deformation. To effectively alleviate this problem, we propose a simple yet effective…

Cited by 0SourceScholar
2024

Visual-Linguistic Representation Learning with Deep Cross-Modality Fusion for Referring Multi-Object Tracking

ICASSP 2024accepted

Referring multi-object tracking is a new rising research topic that aims at detecting and tracking the referred objects in a video sequence based on a natural language expression. Compared with traditional multi-object tracking, this setting guides object tracking with high-level semantic informatio…

Cited by 0SourceScholar
2023

ERBNet: An Effective Representation Based Network for Unbiased Scene Graph Generation

ICASSP 2023accepted

The scene graph generation (SGG) task has attracted increasing attention in recent years. The goal of SGG is to predict relations between pairs of objects within an image. Due to the long-tailed distribution of the dataset annotations, the performance of SGG is still far from satisfactory. To addres…

Cited by 0SourceScholar
2023

Label-Noise Learning with Intrinsically Long-Tailed Data

ICCV 2023poster

Label noise is one of the key factors that lead to the poor generalization of deep learning models. Existing label-noise learning methods usually assume that the ground-truth classes of the training data are balanced. However, the real-world data is often imbalanced, leading to the inconsistency bet…

Cited by 25PDFcodeScholar
2023

Learning to Reconnect Interrupted Trajectories for Weakly Supervised Multi-Object Tracking

ICASSP 2023accepted

Recently, some weakly supervised multi-object tracking (MOT) methods learn identity embedding features with pseudo identity labels rather than the high-cost manual ones. However, these pseudo identity labels may contain many false or missing identities, which adversely affect the optimization of tra…

Cited by 0SourceScholar
2023

Long-Tailed Visual Recognition via Self-Heterogeneous Integration With Knowledge Excavation

CVPR 2023poster

Deep neural networks have made huge progress in the last few decades. However, as the real-world data often exhibits a long-tailed distribution, vanilla deep models tend to be heavily biased toward the majority classes. To address this problem, state-of-the-art methods usually adopt a mixture of exp…

2023

MRCN: A Novel Modality Restitution and Compensation Network for Visible-Infrared Person Re-identification

AAAI 2023technical

Visible-infrared person re-identification (VI-ReID), which aims to search identities across different spectra, is a challenging task due to large cross-modality discrepancy between visible and infrared images. The key to reduce the discrepancy is to filter out identity-irrelevant interference and ef…

Cited by 43SourcePDFScholar
2023

Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation

ICASSP 2023accepted

Personalized Federated Learning (PFL) aims to learn personalized models for each client based on the knowledge across all clients in a privacy-preserving manner. Existing PFL methods generally assume that the underlying global data across all clients are uniformly distributed without considering the…

Cited by 0SourceScholar
2022

A New Framework for Multiple Deep Correlation Filters Based Object Tracking

ICASSP 2022accepted

In recent years, Correlation Filter (CF) based tracking methods using Convolutional Neural Network (CNN) features have achieved the state-of-the-art performance for object tracking. However, how to design an efficient deep CF based tracking method has not been well studied in the literature. To addr…

Cited by 0SourceScholar
2022

Bounding Box Distribution Learning and Center Point Calibration for Robust Visual Tracking

ICASSP 2022accepted

Visual tracking aims at both robust target classification and accurate localization. However, the reliability of the target bounding box and classification score are not properly addressed by most existing trackers, resulting in inaccurate tracking performance. In this paper, we propose to learn bou…

Cited by 0SourceScholar
2022

Federated Learning on Heterogeneous and Long-Tailed Data via Classifier Re-Training with Federated Features

IJCAI 2022poster

Federated learning (FL) provides a privacy-preserving solution for distributed machine learning tasks. One challenging problem that severely damages the performance of FL models is the co-occurrence of data heterogeneity and long-tail distribution, which frequently appears in real FL applications. I…

2022

Learn-to-Decompose: Cascaded Decomposition Network for Cross-Domain Few-Shot Facial Expression Recognition

ECCV 2022poster

"Most existing compound facial expression recognition (FER) methods rely on large-scale labeled compound expression data for training. However, collecting such data is labor-intensive and time-consuming. In this paper, we address the compound FER task in the cross-domain few-shot learning (FSL) sett…

2022

Multi-Focus Guided Semantic Aggregation for Video Object Detection

ICASSP 2022accepted

For the task of video object detection, it is useful to aggregate semantic information from supporting frames. However, existing methods only focus on the current frame during the semantic aggregation, called Single-Focus methods. They neglect semantic information among supporting frames and deterio…

Cited by 0SourceScholar
2022

When Facial Expression Recognition Meets Few-Shot Learning: A Joint and Alternate Learning Framework

AAAI 2022technical

Human emotions involve basic and compound facial expressions. However, current research on facial expression recognition (FER) mainly focuses on basic expressions, and thus fails to address the diversity of human emotions in practical scenarios. Meanwhile, existing work on compound FER relies heavil…

Cited by 18SourcePDFScholar
2021

Feature Decomposition and Reconstruction Learning for Effective Facial Expression Recognition

CVPR 2021poster

In this paper, we propose a novel Feature Decomposition and Reconstruction Learning (FDRL) method for effective facial expression recognition. We view the expression information as the combination of the shared information (expression similarities) across different expressions and the unique informa…

Cited by 218PDFScholar
2021

Learning Spatial-Semantic Relationship for Facial Attribute Recognition With Limited Labeled Data

CVPR 2021poster

Recent advances in deep learning have demonstrated excellent results for Facial Attribute Recognition (FAR), typically trained with large-scale labeled data. However, in many real-world FAR applications, only limited labeled data are available, leading to remarkable deterioration in performance for…

Cited by 41PDFScholar