← Search

Aaron Grattafiori

2 accepted papers

2025

Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

ICML 2025poster

Red teaming aims to assess how large language models (LLMs) can produce content that violates norms, policies, and rules set forth during their safety training. However, most existing automated methods in literature are not representative of the way common users exploit the multi-turn conversational…

Cited by 8SourcePDFScholar
2025

WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

NeurIPS 2025poster

Autonomous UI agents powered by AI have tremendous potential to boost human productivity by automating routine tasks such as filing taxes and paying bills. However, a major challenge in unlocking their full potential is security, which is exacerbated by the agent's ability to take action on their us…

Cited by 0SourcecodeScholar