← Search

Zhaoyu Chen

35 accepted papers

2026

Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models

CVPR 2026

Multimodal biomedical Vision-Language Models (VLMs) exhibit immense potential in the field of Continual Learning (CL). However, they confront a core dilemma: how to preserve fine-grained intra-modality features while bridging the significant domain gap across different modalities. To address this ch

Cited by 0SourcecodeScholar
2026

LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops

ICLR 2026poster

Multimodal Large Language Models (MLLMs) have shown great promise but require substantial computational resources during inference. Attackers can exploit this by inducing excessive output, leading to resource exhaustion and service degradation. Prior energy-latency attacks aim to increase generation…

Cited by 0SourceScholar
2026

Secret-Protected Evolution for Differentially Private Synthetic Text Generation

ICLR 2026poster

Text data has become extremely valuable on large language models (LLMs) and even lead to general artificial intelligence (AGI). A lot of high-quality text in the real world is private and cannot be freely used due to privacy concerns. Therefore, differentially private (DP) synthetic text generation…

Cited by 0SourceScholar
2026

Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

AAAI 2026technical

Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the prin

Cited by 0SourcePDFScholar
2026

Unified Multimodal Visual Tracking with Dual Mixture-of-Experts

ICML 2026poster

Multimodal visual object tracking can be divided into to several kinds of tasks (e.g. RGB and RGB+X tracking), based on the input modality. Existing methods often train separate models for each modality or rely on pretrained models to adapt to new modalities, which limits efficiency, scalability, an…

Cited by 0SourceScholar
2025

Boosting Adversarial Transferability with Spatial Adversarial Alignment

NeurIPS 2025poster

Deep neural networks are vulnerable to adversarial examples that exhibit transferability across various models. Numerous approaches are proposed to enhance the transferability of adversarial examples, including advanced optimization, data augmentation, and model modifications. However, these methods…

Cited by 0SourceScholar
2025

Debiased Multimodal Understanding for Human Language Sequences

AAAI 2025technical

Human multimodal language understanding (MLU) is an indispensable component of expression analysis (e.g., sentiment or humor) from heterogeneous modalities, including visual postures, linguistic contents, and acoustic behaviours. Existing works invariably focus on designing sophisticated structures…

Cited by 1SourcePDFScholar
2025

Dynamic Semantic-Aware Correlation Modeling for UAV Tracking

NeurIPS 2025poster

UAV tracking can be widely applied in scenarios such as disaster rescue, environmental monitoring, and logistics transportation. However, existing UAV tracking methods predominantly emphasize speed and lack exploration in semantic awareness, which hinders the search region from extracting accurate l…

Cited by 0SourceScholar
2025

Enhancing Diffusion-based Unrestricted Adversarial Attacks via Adversary Preferences Alignment

NeurIPS 2025poster

Preference alignment in diffusion models has primarily focused on benign human preferences (e.g., aesthetic). In this paper, we propose a novel perspective: framing unrestricted adversarial example generation as a problem of aligning with adversary preferences. Unlike benign alignment, adversarial a…

Cited by 0SourceScholar
2025

General Compression Framework for Efficient Transformer Object Tracking

ICCV 2025poster

Previous works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex trainin…

2025

Improving Factuality in Large Language Models via Decoding-Time Hallucinatory and Truthful Comparators

AAAI 2025technical

Despite their remarkable capabilities, Large Language Models (LLMs) are prone to generate responses that contradict verifiable facts, i.e., unfaithful hallucination content. Existing efforts generally focus on optimizing model parameters or editing semantic representations, which compromise the inte…

2025

Joint-Wise Distributed Perception Graph Convolutional Network for Skeleton-Based Action Recognition

ICASSP 2025accepted

Recent studies have achieved remarkable results for action recognition with human skeletal data by utilizing graph convolutional models. Traditional approaches typically aggregate local spatio-temporal information bottom-up to form a single spatio-temporal global understanding. However, this method…

Cited by 0SourceScholar
2025

KAN-HyperpointNet for Point Cloud Sequence-Based 3D Human Action Recognition

ICASSP 2025accepted

Point cloud sequence-based 3D action recognition has achieved impressive performance and efficiency. However, existing point cloud sequence modeling methods cannot adequately balance the precision of limb micro-movements with the integrity of posture macro-structure, leading to the loss of crucial i…

Cited by 0SourceScholar
2025

Pruning for Sparse Diffusion Models Based on Gradient Flow

ICASSP 2025accepted

Diffusion Models (DMs) have impressive capabilities among generation models, but are limited to slower inference speeds and higher computational costs. Previous works utilize one-shot structure pruning to derive lightweight DMs from pre-trained ones, but this approach often leads to a significant dr…

Cited by 0SourceScholar
2025

RPPFL: Robust and Privacy-Preserving Federated Learning via Trusted Execution Environments

ICASSP 2025accepted

Federated Learning (FL) is a distributed framework that enables multi-participant collaborative model training without the need for data sharing. Despite its advantages, FL is vulnerable to poisoning and inference attacks, which compromise model accuracy and data privacy. Trusted execution environme…

Cited by 0SourceScholar
2025

Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection

ICCV 2025poster

Pre-trained vision-language models have exhibited remarkable abilities in detecting out-of-distribution (OOD) samples. However, some challenging OOD samples, which lie close to in-distribution (InD) data in image feature space, can still lead to misclassification. The emergence of foundation models…

2025

dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis

CVPR 2025poster

Federated learning has wide applications in the medical field. It enables knowledge sharing among different healthcare institutes while protecting patients' privacy. However, existing federated learning systems are typically centralized, requiring clients to upload client-specific knowledge to a cen…

Cited by 0SourcePDFScholar
2024

De-confounded Data-free Knowledge Distillation for Handling Distribution Shifts

CVPR 2024poster

Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However a long-overlooked iss…

Cited by 6SourcePDFScholar
2024

FAMIM: A Novel Frequency-Domain Augmentation Masked Image Model Framework for Domain Generalizable Face Anti-Spoofing

ICASSP 2024accepted

While existing face anti-spoofing (FAS) methods have achieved high performance on in-domain datasets, good generalization is crucial for their real-world application. Previous domain generalizable FAS methods have attempted to identify common features of live samples from different domains in the sp…

Cited by 0SourceScholar
2024

OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

CVPR 2024highlight

Visual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility tracking tasks can be divided into RGB tracking and RGB+X (e.g. RGB+N and RGB+D) tracking. Despite the different input modalities the core aspec…

Cited by 62SourcePDFScholar
2024

Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation

AAAI 2024technical

Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impract…

Cited by 9SourcePDFScholar
2024

Self-Cooperation Knowledge Distillation for Novel Class Discovery

ECCV 2024poster

"Novel Class Discovery (NCD) aims to discover unknown and novel classes in an unlabeled set by leveraging knowledge already learned about known classes. Existing works focus on instance-level or class-level knowledge representation and build a shared representation space to achieve performance impro…

Cited by 4SourcePDFScholar
2024

Towards Multimodal Sentiment Analysis Debiasing via Bias Purification

ECCV 2024poster

"Multimodal Sentiment Analysis (MSA) aims to understand human intentions by integrating emotion-related clues from diverse modalities, such as visual, language, and audio. Unfortunately, the current MSA task invariably suffers from unplanned dataset biases, particularly multimodal utterance-level la…

Cited by 18SourcePDFScholar
2023

AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving Perception

ICCV 2023poster

Driver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIs…

Cited by 54PDFcodeScholar
2023

Adversarial Contrastive Distillation with Adaptive Denoising

ICASSP 2023accepted

Adversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not a…

Cited by 0SourceScholar
2023

Content-based Unrestricted Adversarial Attack

NeurIPS 2023poster

Unrestricted adversarial attacks typically manipulate the semantic content of an image (e.g., color or texture) to create adversarial examples that are both effective and photorealistic, demonstrating their ability to deceive human perception and deep neural networks with stealth and success. Howeve…

Cited by 86SourcePDFScholar
2023

Context De-Confounded Emotion Recognition

CVPR 2023poster

Context-Aware Emotion Recognition (CAER) is a crucial and challenging task that aims to perceive the emotional states of the target person with contextual information. Recent approaches invariably focus on designing sophisticated architectures or mechanisms to extract seemingly meaningful representa…

2023

Efficient Decision-based Black-box Patch Attacks on Video Recognition

ICCV 2023poster

Although Deep Neural Networks (DNNs) have demonstrated excellent performance, they are vulnerable to adversarial patches that introduce perceptible and localized perturbations to the input. Generating adversarial patches on images has received much attention, while adversarial patches on videos have…

Cited by 23PDFScholar
2023

Explicit and Implicit Knowledge Distillation via Unlabeled Data

ICASSP 2023accepted

Data-free knowledge distillation is a challenging model lightweight task for scenarios in which the original dataset is not available. Previous methods require a lot of extra computational costs to update one or more generators and their naive imitate-learning lead to lower distillation efficiency.…

Cited by 0SourceScholar
2023

Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation

ICCV 2023poster

Learning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficie…

Cited by 25PDFScholar
2023

LVOS: A Benchmark for Long-term Video Object Segmentation

ICCV 2023poster

Existing video object segmentation (VOS) benchmarks focus on short-term videos which just last about 3-5 seconds and where objects are visible most of the time. These videos are poorly representative of practical applications, and the absence of long-term datasets restricts further investigation of…

Cited by 60PDFcodeScholar
2022

CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes

AAAI 2022technical

Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models,…

2022

Efficient Universal Shuffle Attack for Visual Object Tracking

ICASSP 2022accepted

Recently, adversarial attacks have been applied in visual object tracking to deceive deep trackers by injecting imperceptible perturbations into video frames. However, previous work only generates the video-specific perturbations, which restricts its application scenarios. In addition, existing atta…

Cited by 0SourceScholar
2022

Towards Practical Certifiable Patch Defense With Vision Transformer

CVPR 2022poster

Patch attacks, one of the most threatening forms of physical attack in adversarial examples, can lead networks to induce misclassification by modifying pixels arbitrarily in a continuous region. Certifiable patch defense can guarantee robustness that the classifier is not affected by patch attacks.…

Cited by 81PDFScholar