← Search

Wenbo Jiang

15 accepted papers

2026

CBV: Clean-label Backdoor Attacks on Vision Language Models via Diffusion Models

ICML 2026poster

Vision-Language Models (VLMs) have achieved remarkable success in tasks such as image captioning and visual question answering (VQA). However, as their applications become increasingly widespread, recent studies have revealed that VLMs are vulnerable to backdoor attacks. Existing backdoor attacks on…

Cited by 0SourceScholar
2026

DIVER: Diving Deeper into Distilled Data via Expressive Semantic Recovery

ICML 2026poster

Dataset distillation aims to synthesize a compact proxy dataset that is unreadable or non-raw from the original dataset for privacy protection and highly efficient learning. However, previous approaches typically adopt a single-stage distillation paradigm, which suffers from learning specific patter…

Cited by 0SourceScholar
2026

Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models

AAAI 2026technical

Large language models (LLMs) increasingly support multilingual understanding and generation. Meanwhile, efforts to interpret their internal mechanisms have emerged, offering insights to enhance multilingual performance. While multi-head self-attention (MHA) has proven critical in many areas, its rol

Cited by 0SourcePDFScholar
2026

MPMA: Preference Manipulation Attack Against Model Context Protocol

AAAI 2026technical

Model Context Protocol (MCP) standardizes interface mapping for large language models (LLMs) to access external data and tools, which revolutionizes the paradigm of tool selection and facilitates the rapid expansion of the LLM agent tool ecosystem. However, as the MCP is increasingly adopted, third-

Cited by 0SourcePDFScholar
2026

Semantic-level Backdoor Attack against Text-to-Image Diffusion Models

ICML 2026poster

Text-to-image (T2I) diffusion models are widely adopted for their strong generative capabilities, yet remain vulnerable to backdoor attacks. Existing attacks typically rely on fixed textual triggers and single-entity backdoor targets, making them highly susceptible to enumeration-based input defense…

Cited by 0SourceScholar
2026

TEAR: Temporal-aware Automated Red-teaming for Text-to-Video Models

CVPR 2026

Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods, which focus on static image and text generation, are insufficient to c

Cited by 0SourceScholar
2025

BadRefSR: Backdoor Attacks Against Reference-based Image Super Resolution

ICASSP 2025accepted

Reference-based image super-resolution (RefSR) represents a promising advancement in super-resolution (SR). In contrast to single-image super-resolution (SISR), RefSR leverages an additional reference image to help recover high-frequency details, yet its vulnerability to backdoor attacks has not bee…

Cited by 0SourceScholar
2025

Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study

EMNLP 2025

Large Audio-Language Models (LALMs) are increasingly deployed in real-world applications, yet their robustness against malicious audio injection remains underexplored. To address this gap, this study systematically evaluates five leading LALMs across four attack scenarios: Audio Interference Attack,

Cited by 0SourcePDFScholar
2025

Omni-Angle Assault: An Invisible and Powerful Physical Adversarial Attack on Face Recognition

ICML 2025poster

Deep learning models employed in face recognition (FR) systems have been shown to be vulnerable to physical adversarial attacks through various modalities, including patches, projections, and infrared radiation. However, existing adversarial examples targeting FR systems often suffer from issues suc…

Cited by 0SourcePDFScholar
2025

PRESS: Defending Privacy in Retrieval-Augmented Generation via Embedding Space Shifting

ICASSP 2025accepted

Retrieval-augmented generation (RAG) expands the capabilities of large language models (LLMs) in various applications by integrating relevant information retrieved from external data sources. However, the RAG systems are exposed to substantial privacy risks during the information retrieval process,…

Cited by 0SourceScholar
2025

The Fluorescent Veil: A Stealthy and Effective Physical Adversarial Patch Against Traffic Sign Recognition

NeurIPS 2025poster

Recently, traffic sign recognition (TSR) systems have become a prominent target for physical adversarial attacks. These attacks typically rely on conspicuous stickers and projections, or using invisible light and acoustic signals that can be easily blocked. In this paper, we introduce a novel attack…

Cited by 0SourceScholar
2025

The Ripple Effect: On Unforeseen Complications of Backdoor Attacks

ICML 2025poster

Recent research highlights concerns about the trustworthiness of third-party Pre-Trained Language Models (PTLMs) due to potential backdoor attacks. These backdoored PTLMs, however, are effective only for specific pre-defined downstream tasks. In reality, these PTLMs can be adapted to many other unre…

2025

Watch Out for Your Guidance on Generation! Exploring Conditional Backdoor Attacks against Large Language Models

AAAI 2025technical

Mainstream backdoor attacks on large language models (LLMs) typically set a fixed trigger in the input instance and specific responses for triggered queries. However, the fixed trigger setting (e.g., unusual words) may be easily detected by human detection, limiting the effectiveness and practicali…

Cited by 0SourcePDFScholar