2026
RedDebate: Safer Responses Through Multi-Agent Red Teaming Debates
ICML 2026poster
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their own unsafe behaviors. Existing AI safety approaches often rely on costly human evaluation or isolated single-model assessment, both constrained by…