← Search

Yingying Chen

28 accepted papers

2026

AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection

AAAI 2026technical

Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performanc

Cited by 0SourcePDFScholar
2026

Benchmarking the Scientific Mind: Toward Evaluation of Complex-Reasoning Biomedical VQA

ICML 2026poster

Despite progress of Multimodal Large Language Models (MLLMs) in biomedical visual question answering (VQA), existing benchmarks provide limited assessment of their scientific reasoning capabilities. Most datasets adopt single-image question construction and outcome-oriented evaluation, where correct…

Cited by 0SourceScholar
2026

Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection

AAAI 2026technical

Despite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce

Cited by 0SourcePDFScholar
2026

Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning

CVPR 2026

Recent advances in audio-visual representation learning have shown the value of combining contrastive alignment with masked reconstruction. However, jointly optimizing these objectives in a single forward pass forces the contrastive branch to rely on randomly visible patches designed for reconstruct

Cited by 0SourceScholar
2025

FLARE: A Framework for Stellar Flare Forecasting Using Stellar Physical Properties and Historical Records

IJCAI 2025

Stellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecast

2025

Fine-grained Vital Sign Reconstruction through Machine Learning on Multi-channel Radar Signals

ICASSP 2025accepted

Monitoring vital signs such as breathing rate (BR) and heart rate (HR) is crucial for early detection of health issues and supports a wide range of health-related applications. Traditional monitoring methods often involve body-attached medical devices, which can be intrusive and inconvenient for con…

Cited by 0SourceScholar
2025

LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing

ICASSP 2025accepted

Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often lack alignment, thereby introducing extra noise during modal…

Cited by 0SourceScholar
2025

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

ICCV 2025poster

The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the limitations of the weakly-supervised and the deficiencies of the model architecture, existing methods are lacking in sim…

2025

UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection

CVPR 2025poster

Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with sp…

2024

AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models

AAAI 2024technical

Large Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific d…

2024

Clean & Compact: Efficient Data-Free Backdoor Defense with Model Compactness

ECCV 2024poster

"Deep neural networks (DNNs) have been widely deployed in real-world, mission-critical applications, necessitating effective approaches to protect deep learning models against malicious attacks. Motivated by the high stealthiness and potential harm of backdoor attacks, a series of backdoor defense m…

Cited by 2SourcePDFScholar
2023

Benchmarking and Analyzing Robust Point Cloud Recognition: Bag of Tricks for Defending Adversarial Examples

ICCV 2023poster

Deep Neural Networks (DNNs) for 3D point cloud recognition are vulnerable to adversarial examples, threatening their practical deployment. Despite the many research endeavors have been made to tackle this issue in recent years, the diversity of adversarial examples on 3D point clouds makes them more…

Cited by 5PDFcodeScholar
2023

DynGMP: Graph Neural Network-Based Motion Planning in Unpredictable Dynamic Environments

IROS 2023poster

Neural networks have already demonstrated attractive performance for solving motion planning problems, especially in static and predictable environments. However, efficient neural planners that can adapt to unpredictable dynamic environments, a highly demanded scenario in many practical applications…

Cited by 2SourceScholar
2022

C2AM Loss: Chasing a Better Decision Boundary for Long-Tail Object Detection

CVPR 2022poster

Long-tail object detection suffers from poor performance on tail categories. We reveal that the real culprit lies in the extremely imbalanced distribution of the classifier's weight norm. For conventional softmax cross-entropy loss, such imbalanced weight norm distribution yields ill conditioned dec…

Cited by 28PDFScholar
2022

Forecasting Asset Dependencies to Reduce Portfolio Risk

AAAI 2022technical

Financial assets exhibit dependence structures, i.e., movements of their prices or returns show various correlations. Knowledge of assets’ price dependencies can help investors to create a diversified portfolio, aiming to reduce portfolio risk due to the high volatility of the financial market. Sinc…

Cited by 9SourcePDFScholar
2022

Invisible and Efficient Backdoor Attacks for Compressed Deep Neural Networks

ICASSP 2022accepted

Compressed deep neural network (DNN) models have been widely deployed in many resource-constrained platforms and devices. However, the security issue of the compressed models, especially their vulnerability against backdoor attacks, is not well explored yet. In this paper, we study the feasibility o…

Cited by 0SourceScholar
2022

RIBAC: Towards Robust and Imperceptible Backdoor Attack against Compact DNN

ECCV 2022poster

"Recently backdoor attack has become an emerging threat to the security of deep neural network (DNN) models. To date, most of the existing studies focus on backdoor attack against the uncompressed model; while the vulnerability of compressed DNNs, which are widely used in the practical applications,…

2022

Regularizing Vector Embedding in Bottom-Up Human Pose Estimation

ECCV 2022poster

"The embedding-based method such as Associative Embedding is popular in bottom-up human pose estimation. Methods under this framework group candidate keypoints according to the predicted identity embeddings. However, the identity embeddings of different instances are likely to be linearly inseparabl…

2022

UniVIP: A Unified Framework for Self-Supervised Visual Pre-Training

CVPR 2022poster

Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of…

Cited by 41PDFScholar
2021

Enabling Fast and Universal Audio Adversarial Attack Using Generative Model

AAAI 2021technical

Recently, the vulnerability of deep neural network (DNN)-based audio systems to adversarial attacks has obtained increasing attention. However, the existing audio adversarial attacks allow the adversary to possess the entire user's audio input as well as granting sufficient time budget to generate t…

Cited by 79SourcePDFScholar
2021

Improving Multiple Object Tracking With Single Object Tracking

CVPR 2021poster

Despite considerable similarities between multiple object tracking (MOT) and single object tracking (SOT) tasks, modern MOT methods have not benefited from the development of SOT ones to achieve satisfactory performance. The major reason for this situation is that it is inappropriate and inefficient…

Cited by 148PDFScholar
2020

BDD100K: A Diverse Driving Dataset for Heterogeneous Multitask Learning

CVPR 2020oral

Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving. Researchers are usually constrained to study a small set of problems on one dataset, while real-world computer vision appl…

Cited by 2855PDFScholar
2020

Learning Feature Embeddings for Discriminant Model based Tracking

ECCV 2020poster

After observing that the features used in most online discriminatively trained trackers are not optimal, in this paper, we propose a novel and effective architecture to learn optimal feature embeddings for online discriminative tracking. Our method, called DCFST, integrates the solver of a discrimin…

Cited by 111SourcePDFScholar
2020

Occlusion-Aware Siamese Network for Human Pose Estimation

ECCV 2020poster

Pose estimation usually suffers from varying degrees of performance degeneration owing to occlusion. To conquer this dilemma, we propose an occlusion-aware siamese network to improve the performance. Specifically, we introduce scheme of feature erasing and reconstruction. Firstly, we utilize attenti…

Cited by 50SourcePDFScholar
2020

Real-Time, Universal, and Robust Adversarial Attacks Against Speaker Recognition Systems

ICASSP 2020accepted

As the popularity of voice user interface (VUI) exploded in recent years, speaker recognition system has emerged as an important medium of identifying a speaker in many security-required applications and services. In this paper, we propose the first real-time, universal, and robust adversarial attac…

Cited by 0SourceScholar