2025
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
ICLR 2025oral
Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policymakers, model providers, and researchers in the AI and cybersecurity communities are interested in quantifying the capabil…