2025
SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks
NeurIPS 2025poster
Rigorous security-focused evaluation of large language model (LLM) agents is imperative for establishing trust in their safe deployment throughout the software development lifecycle. However, existing benchmarks largely rely on synthetic challenges or simplified vulnerability datasets that fail to c…