← Search

Tanqiu Jiang

5 accepted papers

2026

Dynamic Token Reweighting for Robust Vision-Language Models

CVPR 2026

Large vision-language models (VLMs) are highly vulnerable to multimodal jailbreak attacks that exploit visual-textual interactions to bypass safety guardrails. In this paper, we present DTR, a novel inference-time defense that mitigates multimodal jailbreak attacks through optimizing the model's key

Cited by 0SourcecodeScholar
2025

RAPID: Retrieval Augmented Training of Differentially Private Diffusion Models

ICLR 2025poster

Differentially private diffusion models (DPDMs) harness the remarkable generative capabilities of diffusion models while enforcing differential privacy (DP) for sensitive data. However, existing DPDM training approaches often suffer from significant utility loss, large memory footprint, and expensiv…

2025

RobustKV: Defending Large Language Models against Jailbreak Attacks via KV Eviction

ICLR 2025poster

Jailbreak attacks circumvent LLMs' built-in safeguards by concealing harmful queries within adversarial prompts. While most existing defenses attempt to mitigate the effects of adversarial prompts, they often prove inadequate as adversarial prompts can take arbitrary, adaptive forms. This paper intr…

Cited by 4SourcePDFScholar