← Search

Shengfang ZHAI

8 accepted papers

2026

*MemPot*: Defend Against Memory Extraction Attack with Optimized Honeypots

ICML 2026poster

Large Language Model (LLM)-based agents employ external and internal memory systems to handle complex, goal-oriented tasks, yet this exposes them to severe extraction attacks, and corresponding defenses are currently lacking. In this paper, we propose *MemPot*, the first theoretically verified defen…

Cited by 0SourceScholar
2026

Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems

ICLR 2026poster

Retrieval-Augmented Generation (RAG) systems enhance large language models (LLMs) by incorporating external knowledge bases, but this may expose them to extraction attacks, leading to potential copyright and privacy risks. However, existing extraction methods typically rely on malicious inputs such…

Cited by 0SourceScholar
2025

Efficient Input-level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation Variation

ICCV 2025poster

In recent years, text-to-image (T2I) diffusion models have gained significant attention for their ability to generate high-quality images reflecting text prompts. However, their growing popularity has also led to the emergence of backdoor threats, posing substantial risks. Currently, effective defen…

Cited by 0SourcePDFScholar
2025

GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning

NeurIPS 2025poster

To enhance the safety of VLMs, this paper introduces a novel reasoning-based VLM guard model dubbed GuardReasoner-VL. The core idea is to incentivize the guard model to deliberatively reason before making moderation decisions via online RL. First, we construct GuardReasoner-VLTrain, a reasoning corp…

Cited by 0SourcecodeScholar
2024

Membership Inference on Text-to-Image Diffusion Models via Conditional Likelihood Discrepancy

NeurIPS 2024poster

Text-to-image diffusion models have achieved tremendous success in the field of controllable image generation, while also coming along with issues of privacy leakage and data copyrights. Membership inference arises in these contexts as a potential auditing method for detecting unauthorized data usag…

2024

Security Equivalence Assessment between Cloud Standards by Mapping of Control Items

ICASSP 2024accepted

The rise of new industries, such as the Internet of Things and Smart Healthcare, has brought many cross-cloud business opportunities for cloud computing and posed new challenges to the cloud security. Traditionally, security can be assessed by compliance checking when selecting cloud services. Howev…

Cited by 0SourceScholar
2024

TRLS: A Time Series Representation Learning Framework Via Spectrogram for Medical Signal Processing

ICASSP 2024accepted

Representation learning frameworks in unlabeled time series have been proposed for medical signal processing. Despite the numerous excellent progresses have been made in previous works, we observe the representation extracted for the time series still does not generalize well. In this paper, we pres…

Cited by 0SourceScholar
2023

NCL: Textual Backdoor Defense Using Noise-Augmented Contrastive Learning

ICASSP 2023accepted

At present, backdoor attacks attract attention as they do great harm to deep learning models. By poisoning the training data, the adversary makes the model trained based on this dataset being injected with a backdoor. In the field of text, however, existing works do not provide sufficient defense ag…

Cited by 0SourceScholar