← Search

Avidan Shah

2 accepted papers

2026

Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework

ICML 2026spotlight

The Rapid Response (RR) framework (Peng et al., 2024), deployed in production systems including Anthropic’s ASL-3 safeguards (Anthropic, 2025), dynamically adapts jailbreak detection classifiers by generating synthetic training data from emerging attacks. We reveal that prompt injection can infiltra…

Cited by 0SourceScholar
2025

Stronger Universal and Transferable Attacks by Suppressing Refusals

NAACL 2025long

Making large language models (LLMs) safe for mass deployment is a complex and ongoing challenge. Efforts have focused on aligning models to human preferences (RLHF), essentially embedding a “safety feature” into the model’s parameters. The Greedy Coordinate Gradient (GCG) algorithm (Zou et al., 2023…

Cited by 0SourcePDFScholar