2026
PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities
ICLR 2026poster
The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PAC…