← Search

Jaewon Chang

2 accepted papers

2026

Rapid Poison: Practical Poisoning Attacks Against the Rapid Response Framework

ICML 2026spotlight

The Rapid Response (RR) framework (Peng et al., 2024), deployed in production systems including Anthropic’s ASL-3 safeguards (Anthropic, 2025), dynamically adapts jailbreak detection classifiers by generating synthetic training data from emerging attacks. We reveal that prompt injection can infiltra…

Cited by 0SourceScholar
2024

PubDef: Defending Against Transfer Attacks From Public Models

ICLR 2024poster

Adversarial attacks have been a looming and unaddressed threat in the industry. However, through a decade-long history of the robustness evaluation literature, we have learned that mounting a strong or optimal attack is challenging. It requires both machine learning and domain expertise. In other wo…