← Search

Maya Pavlova

3 accepted papers

2025

AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents

NeurIPS 2025poster

Autonomous AI agents that can follow instructions and perform complex multi-step tasks have tremendous potential to boost human productivity. However, to perform many of these tasks, the agents need access to personal information from their users, raising the question of whether they are capable of…

Cited by 0SourcecodeScholar
2025

Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

ICML 2025poster

Red teaming aims to assess how large language models (LLMs) can produce content that violates norms, policies, and rules set forth during their safety training. However, most existing automated methods in literature are not representative of the way common users exploit the multi-turn conversational…

Cited by 8SourcePDFScholar