← Search

Yechao Zhang

9 accepted papers

2026

Dual-View Inference Attack: Machine Unlearning Amplifies Privacy Exposure

AAAI 2026technical

Machine unlearning is a newly popularized technique for removing specific training data from a trained model, enabling it to comply with data deletion requests. While it protects the rights of users requesting unlearning, it also introduces new privacy risks. Prior works have primarily focused on th

Cited by 0SourcePDFScholar
2026

MHB: Medical Hallucination Benchmark for Large Language Models in Complex Clinical Tasks

AAAI 2026technical

The integration of Large Language Models (LLMs) into clinical applications presents transformative potential but is undermined by the critical risk of hallucination, the generation of plausible but factually incorrect information. Such failures pose a direct threat to patient safety and the integrit

Cited by 0SourcePDFScholar
2026

VideoSEAL: Separating Planning from Answer Authority for Agentic Long Video Understanding

ICML 2026poster

Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long videos introduce long-horizon search and verification, which often necessitates multi-turn, agentic interaction. We show t…

Cited by 0SourceScholar
2025

Improving Generalization of Universal Adversarial Perturbation via Dynamic Maximin Optimization

AAAI 2025technical

Deep neural networks (DNNs) are susceptible to universal adversarial perturbations (UAPs). These perturbations are meticulously designed to fool the target model universally across all sample classes. Unlike instance-specific adversarial examples (AEs), generating UAPs is more complex because they m…

2025

MARS: A Malignity-Aware Backdoor Defense in Federated Learning

NeurIPS 2025poster

Federated Learning (FL) is a distributed paradigm aimed at protecting participant data privacy by exchanging model parameters to achieve high-quality model training. However, this distributed nature also makes FL highly vulnerable to backdoor attacks. Notably, the recently proposed state-of-the-art…

Cited by 0SourceScholar
2025

Transferable Direct Prompt Injection via Activation-Guided MCMC Sampling

EMNLP 2025

Direct Prompt Injection (DPI) attacks pose a critical security threat to Large Language Models (LLMs) due to their low barrier of execution and high potential damage. To address the impracticality of existing white-box/gray-box methods and the poor transferability of black-box methods, we propose an

Cited by 0SourcePDFScholar
2024

Unlearnable 3D Point Clouds: Class-wise Transformation Is All You Need

NeurIPS 2024poster

Traditional unlearnable strategies have been proposed to prevent unauthorized users from training on the 2D image data. With more 3D point cloud data containing sensitivity information, unauthorized usage of this new type data has also become a serious concern. To address this, we propose the first…

2022

Protecting Facial Privacy: Generating Adversarial Identity Masks via Style-Robust Makeup Transfer

CVPR 2022poster

While deep face recognition (FR) systems have shown amazing performance in identification and verification, they also arouse privacy concerns for their excessive surveillance on users, especially for public face images widely spread on social networks. Recently, some studies adopt adversarial exampl…

Cited by 130PDFcodeScholar