← Search

Chongyang Zhang

19 accepted papers

2026

MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

CVPR 2026

In the progress of industrial anomaly detection, general anomaly detection (GAD) is an emerging trend and also the ultimate goal. Unlike the conventional single- and multi-class AD, general AD aims to train a general AD model that can directly detect anomalies in diverse novel classes without any re

Cited by 0SourcecodeScholar
2025

ADPretrain: Advancing Industrial Anomaly Detection via Anomaly Representation Pretraining

NeurIPS 2025poster

The current mainstream and state-of-the-art anomaly detection (AD) methods are substantially established on pretrained feature networks yielded by ImageNet pre- training. However, regardless of supervised or self-supervised pretraining, the pretraining process on ImageNet does not match the goal of…

Cited by 0SourcecodeScholar
2025

Beyond Label Semantics: Language-Guided Action Anatomy for Few-shot Action Recognition

ICCV 2025poster

Few-shot action recognition (FSAR) aims to classify human actions in videos with only a small number of labeled samples per category. The scarcity of training data has driven recent efforts to incorporate additional modalities, particularly text. However, the subtle variations in human posture, moti…

Cited by 0SourcePDFScholar
2025

GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration

CVPR 2025poster

GUI agents hold significant potential to enhance the experience and efficiency of human-device interaction. However, current methods face challenges in generalizing across applications (apps) and tasks, primarily due to two fundamental limitations in existing datasets. First, these datasets overlook…

2025

Make Unseen Clear: Occluder Removal for Complete 3D Pedestrian Detection

ICASSP 2025accepted

In autonomous driving, the ability to detect pedestrians accurately is crucial for safety. Some detectors, however, often struggle with occlusions, where pedestrians partially hidden behind objects appear incomplete and are harder to be identified accurately. To alleviate this issue, we introduce Cl…

Cited by 0SourceScholar
2024

ResAD: A Simple Framework for Class Generalizable Anomaly Detection

NeurIPS 2024spotlight

This paper explores the problem of class-generalizable anomaly detection, where the objective is to train one unified AD model that can generalize to detect anomalies in diverse classes from different domains without any retraining or fine-tuning on the target data. Because normal feature representa…

2023

Explicit Boundary Guided Semi-Push-Pull Contrastive Learning for Supervised Anomaly Detection

CVPR 2023poster

Most anomaly detection (AD) models are learned using only normal samples in an unsupervised way, which may result in ambiguous decision boundary and insufficient discriminability. In fact, a few anomaly samples are often available in real-world applications, the valuable knowledge of known anomalies…

2023

Focus the Discrepancy: Intra- and Inter-Correlation Learning for Image Anomaly Detection

ICCV 2023poster

Humans recognize anomalies through two aspects: larger patch-wise representation discrepancies and weaker patch-to-normal-patch correlations. However, the previous AD methods didn't sufficiently combine the two complementary aspects to design AD models. To this end, we find that Transformer can idea…

Cited by 27PDFcodeScholar
2023

Modulation-Based Center Alignment and Motion Mining for Spatial Temporal Action Detection

ICASSP 2023accepted

The goal of spatial-temporal action detection is to generate spatial-temporally aligned action tubes. Most of the existing 2D CNN-based solutions directly aggregate temporal adjacent contexts through frames without alignment. The misaligned spatial-temporal contextual features might lead to chaotic…

Cited by 0SourceScholar
2023

One-for-All: Proposal Masked Cross-Class Anomaly Detection

AAAI 2023technical

One of the most challenges for anomaly detection (AD) is how to learn one unified and generalizable model to adapt to multi-class especially cross-class settings: the model is trained with normal samples from seen classes with the objective to detect anomalies from both seen and unseen classes. In t…

2022

Out-of-Distribution Identification: Let Detector Tell Which I Am Not Sure

ECCV 2022poster

"The superior performance of object detectors is often established under the condition that the test samples are in the same distribution as the training data. However, in most practical applications, out-of-distribution (OOD) instances are inevitable and usually lead to detection uncertainty. In th…

Cited by 9SourcePDFScholar
2021

Embracing Uncertainty: Decoupling and De-Bias for Robust Temporal Grounding

CVPR 2021poster

Temporal grounding aims to localize temporal boundaries within untrimmed videos by language queries, but it faces the challenge of two types of inevitable human uncertainties: query uncertainty and label uncertainty. The two uncertainties stem from human subjectivity, leading to limited generalizati…

Cited by 62PDFScholar
2020

Global and Local Discriminative Patches Exploiting for Action Recognition

ICASSP 2020accepted

Recent human action recognition models mainly focus on exploiting human features, such as pose or skeleton features. However, most of these methods do not pay enough attention to action-related backgrounds. In this work we propose a novel multi-stream features fusion framework based on discriminativ…

Cited by 0SourceScholar
2020

Where, What, Whether: Multi-Modal Learning Meets Pedestrian Detection

CVPR 2020poster

Pedestrian detection benefits greatly from deep convolutional neural networks (CNNs). However, it is inherently hard for CNNs to handle situations in the presence of occlusion and scale variation. In this paper, we propose W^3Net, which attempts to address above challenges by decomposing the pedestr…

Cited by 39PDFScholar
2019

Visual Relationship Recognition via Language and Position Guided Attention

ICASSP 2019accepted

Visual relationship recognition, as a challenging task used to distinguish the interactions between object pairs, has received much attention recently. Considering the fact that most visual relationships are semantic concepts defined by human beings, there are many human knowledge, or priors, hidden…

Cited by 0SourceScholar
2016

Learning discriminative and shareable patches for scene classification

ICASSP 2016accepted

This paper addresses the problem of scene classification and proposes learning discriminative and shareable patches (LDSP) method. The main idea of learning discriminative and shareable patches is to discover patches that exhibit both large between-class dissimilarity (discriminative) and large with…

Cited by 0SourceScholar