← Search

Peter A. Beling

3 accepted papers

2025

From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing

EMNLP 2025

Large Language Models (LLMs) have been explored for automating or enhancing penetration testing tasks, but their effectiveness and reliability across diverse attack phases remain open questions. This study presents a comprehensive evaluation of multiple LLM-based agents, ranging from singular to mod

Cited by 0SourcePDFScholar
2025

GENUINE: Graph Enhanced Multi-level Uncertainty Estimation for Large Language Models

EMNLP 2025

Uncertainty estimation is essential for enhancing the reliability of Large Language Models (LLMs), particularly in high-stakes applications. Existing methods often overlook semantic dependencies, relying on token-level probability measures that fail to capture structural relationships within the gen