← Search

Lige Huang

1 accepted papers

2026

PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities

ICLR 2026poster

The increasing autonomy of Large Language Models (LLMs) necessitates a rigorous evaluation of their potential to aid in cyber offense. Existing benchmarks often lack real-world complexity and are thus unable to accurately assess LLMs' cybersecurity capabilities. To address this gap, we introduce PAC…

Cited by 0SourcecodeScholar