← Search

Mahdi Sabbaghi

1 accepted papers

2025

Adversarial Reasoning at Jailbreaking Time

ICML 2025poster

As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important. Recent advances in standardizing, measuring, and scaling test-time compute suggest new methodologies for optimizing models to achieve high performance on ha…