2025
SECODEPLT: A Unified Benchmark for Evaluating the Security Risks and Capabilities of Code GenAI
NeurIPS 2025poster
Existing benchmarks for evaluating the security risks and capabilities (e.g., vulnerability detection) of code-generating large language models (LLMs) face several key limitations: (1) limited coverage of risk and capabilities; (2) reliance on static evaluation metrics such as LLM judgments or rule-…