← Search

Zhan Qin

19 accepted papers

2026

Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization

AAAI 2026technical

While text embeddings enable efficient semantic processing in LLMs, they remain vulnerable to inversion attacks that reconstruct sensitive original text. However, current defense methods typically treat text embeddings from the feature level independently, ignoring the exploitation of the mutual rel

Cited by 0SourcePDFScholar
2026

Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets

ICLR 2026poster

Large Language Models (LLMs) have seen remarkable advancements, achieving state-of-the-art results in diverse applications. Fine-tuning, an important step for adapting LLMs to specific downstream tasks, typically involves further training on corresponding datasets. However, a fundamental discrepancy…

Cited by 0SourceScholar
2026

JANUS: A Lightweight Framework for Jailbreaking Text-to-Image Models via Distribution Optimization

CVPR 2026

Text-to-image (T2I) models such as Stable Diffusion and DALLE remain susceptible to generating harmful or Not-Safe-For-Work (NSFW) content under jailbreak attacks despite deployed safety filters. Existing jailbreak attacks either rely on proxy-loss optimization instead of the true end-to-end objecti

Cited by 0SourcecodeScholar
2026

MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies

AAAI 2026technical

Large Language Models (LLMs) have exhibited remarkable capabilities but remain vulnerable to jailbreaking attacks, which can elicit harmful content from the models by manipulating the input prompts. Existing black-box jailbreaking techniques primarily rely on static prompts crafted with a single, no

Cited by 0SourcePDFScholar
2026

SpatialJB: How Text Distribution Art Becomes The "Jailbreak Key" for LLM Guardrails

ICML 2026poster

While Large Language Models (LLMs) have achieved remarkable success across diverse tasks, they remain vulnerable to jailbreak attacks, which pose significant risks to their secure deployment. Current safetymechanisms primarily rely on output guardrails to filter harmful outputs, yet these defenses a…

Cited by 0SourceScholar
2025

DELMAN: Dynamic Defense Against Large Language Model Jailbreaking with Model Editing

ACL 2025finding

Large Language Models (LLMs) are widely applied in decision making, but their deployment is threatened by jailbreak attacks, where adversarial users manipulate model behavior to bypass safety measures. Existing defense mechanisms, such as safety fine-tuning and model editing, either require extensiv…

2025

Don’t Say No: Jailbreaking LLM by Suppressing Refusal

ACL 2025finding

Ensuring the safety alignment of Large Language Models (LLMs) is critical for generating responses consistent with human values. However, LLMs remain vulnerable to jailbreaking attacks, where carefully crafted prompts manipulate them into producing toxic content. One category of such attacks reformu…

2025

REFINE: Inversion-Free Backdoor Defense via Model Reprogramming

ICLR 2025poster

Backdoor attacks on deep neural networks (DNNs) have emerged as a significant security threat, allowing adversaries to implant hidden malicious behaviors during the model training phase. Pre-processing-based defense, which is one of the most important defense paradigms, typically focuses on input tr…

Cited by 2SourcePDFScholar
2025

Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack

NeurIPS 2025poster

Knowledge distillation (KD) is a vital technique for deploying deep neural networks (DNNs) on resource-constrained devices by transferring knowledge from large teacher models to lightweight student models. While teacher models from third-party platforms may undergo security verification (e.g., backd…

Cited by 0SourcecodeScholar
2025

Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning

NeurIPS 2025poster

Text-to-image (T2I) diffusion models have achieved impressive image generation quality and are increasingly fine-tuned for personalized applications. However, these models often inherit unsafe behaviors from toxic pretraining data, raising growing safety concerns. While recent safety-driven unlearni…

Cited by 0SourcecodeScholar
2025

WMCopier: Forging Invisible Watermarks on Arbitrary Images

NeurIPS 2025poster

Invisible Image Watermarking is crucial for ensuring content provenance and accountability in generative AI. While Gen-AI providers are increasingly integrating invisible watermarking systems, the robustness of these schemes against forgery attacks remains poorly characterized. This is critical, as…

Cited by 0SourcecodeScholar
2024

Towards Reliable and Efficient Backdoor Trigger Inversion via Decoupling Benign Features

ICLR 2024spotlight

Recent studies revealed that using third-party models may lead to backdoor threats, where adversaries can maliciously manipulate model predictions based on backdoors implanted during model training. Arguably, backdoor trigger inversion (BTI), which generates trigger patterns of given benign samples…

Cited by 31SourcePDFScholar
2023

A Large-Scale Pretrained Deep Model for Phishing URL Detection

ICASSP 2023accepted

Phishing attacks have always been a security issue that has attracted great attention in the cyber security community. Recently, the famous pre-trained models is being used as an anti-phishing solution. However, existing studies either simply transfer models pre-trained on text to phishing detection…

Cited by 0SourceScholar
2023

Certified Minimax Unlearning with Generalization Rates and Deletion Capacity

NeurIPS 2023poster

We study the problem of $(\epsilon,\delta)$-certified machine unlearning for minimax models. Most of the existing works focus on unlearning from standard statistical learning models that have a single variable and their unlearning steps hinge on the direct Hessian-based conventional Newton update. W…

Cited by 23SourcePDFScholar
2023

MUter: Machine Unlearning on Adversarially Trained Models

ICCV 2023poster

Machine unlearning is an emerging task of removing the influence of selected training datapoints from a trained model upon data deletion requests, which echoes the widely enforced data regulations mandating the Right to be Forgotten. Many unlearning methods have been proposed recently, achieving sig…

Cited by 27PDFScholar
2022

Backdoor Defense via Decoupling the Training Process

ICLR 2022poster

Recent studies have revealed that deep neural networks (DNNs) are vulnerable to backdoor attacks, where attackers embed hidden backdoors in the DNN model by poisoning a few training samples. The attacked model behaves normally on benign samples, whereas its prediction will be maliciously changed whe…

2021

Feature Importance-Aware Transferable Adversarial Attacks

ICCV 2021poster

Transferability of adversarial examples is of central importance for attacking an unknown model, which facilitates adversarial attacks in more practical scenarios, e.g., blackbox attacks. Existing transferable attacks tend to craft adversarial examples by indiscriminately distorting features to degr…

Cited by 288PDFcodeScholar
2021

From Local to Global Norm Emergence: Dissolving Self-reinforcing Substructures with Incremental Social Instruments

ICML 2021spotlight

Norm emergence is a process where agents in a multi-agent system establish self-enforcing conformity through repeated interactions. When such interactions are confined to a social topology, several self-reinforcing substructures (SRS) may emerge within the population. This prevents a formation of a…

Cited by 12SourcePDFScholar