2025
CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
NeurIPS 2025poster
Recent advances in Large Language Models (LLMs) have spurred transformative applications in various domains, ranging from open-source to proprietary LLMs. However, jailbreak attacks, which aim to break safety alignment and user compliance by tricking the target LLMs into answering harmful and risky…