2025
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models
AAAI 2025technical
As Large Language Models (LLMs) continue to advance in capability and influence, ensuring their security and preventing harmful outputs has become crucial. A promising approach to address these concerns involves training models to automatically generate adversarial prompts for red teaming. However,…