2025
Diversity Seeking Techniques for Red-Teaming Large Language Models
ICASSP 2025accepted
In this paper, we present new techniques for increasing the diversity of red-teaming prompts generated by automated machine learning-based methods, thereby enabling the discovery of more vulnerabilities in large language models. Using reinforcement learning to train models to output effective prompt…