From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
Large Language Models (LLMs) have been explored for automating or enhancing penetration testing tasks, but their effectiveness and reliability across diverse attack phases remain open questions. This study presents a comprehensive evaluation of multiple LLM-based agents, ranging from singular to mod