← Search

Amelia Hardy

2 accepted papers

2025

ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts

EMNLP 2025

Existing LLM red-teaming approaches prioritize high attack success rate, often resulting in high-perplexity prompts. This focus overlooks low-perplexity attacks that are more difficult to filter, more likely to arise during benign usage, and more impactful as negative downstream training examples. I

Cited by 0SourcePDFScholar
2024

BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

NeurIPS 2024spotlight

AI models are increasingly prevalent in high-stakes environments, necessitating thorough assessment of their capabilities and risks. Benchmarks are popular for measuring these attributes and for comparing model performance, tracking progress, and identifying weaknesses in foundation and non-foundati…

Cited by 19SourcePDFScholar