← Search

Paul Kassianik

3 accepted papers

2026

Capability-Based Scaling Trends for LLM-Based Red-Teaming

ICLR 2026poster

As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering approaches may prove ineffective once red-teaming turns into a \emph{weak-to-strong} problem, where target models surpass…

Cited by 0SourcecodeScholar
2025

Adversarial Reasoning at Jailbreaking Time

ICML 2025poster

As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important. Recent advances in standardizing, measuring, and scaling test-time compute suggest new methodologies for optimizing models to achieve high performance on ha…

2024

Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

NeurIPS 2024poster

While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed *jailbreaks*. In this work, we present *Tree of Attacks with Pruning* (TAP), an automated method for generating jailb…