← Search

Hailey Nguyen

2 accepted papers

2025

Automated Red Teaming with GOAT: the Generative Offensive Agent Tester

ICML 2025poster

Red teaming aims to assess how large language models (LLMs) can produce content that violates norms, policies, and rules set forth during their safety training. However, most existing automated methods in literature are not representative of the way common users exploit the multi-turn conversational…

Cited by 8SourcePDFScholar
2025

Backtracking Improves Generation Safety

ICLR 2025oral

Text generation has a fundamental limitation almost by definition: there is no taking back tokens that have been generated, even when they are clearly problematic. In the context of language model safety, when a partial unsafe generation is produced, language models by their nature tend to happily k…

Cited by 14SourcePDFScholar