2025
BaxBench: Can LLMs Generate Correct and Secure Backends?
ICML 2025spotlight
Automatic program generation has long been a fundamental challenge in computer science. Recent benchmarks have shown that large language models (LLMs) can effectively generate code at the function level, make code edits, and solve algorithmic coding tasks. However, to achieve full automation, LLMs s…