2026
Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks
ICLR 2026poster
Despite recent rapid progress in AI safety, current large language models remain vulnerable to adversarial attacks in multi-turn interaction settings, where attackers strategically adapt their prompts across conversation turns and pose a more critical yet realistic challenge. Existing approaches tha…