2025
Rethinking Verification for LLM Code Generation: From Generation to Testing
NeurIPS 2025poster
Large language models (LLMs) have recently achieved notable success in code‑generation benchmarks such as HumanEval and LiveCodeBench. However, a detailed examination reveals that these evaluation suites often comprise only a limited number of homogeneous test cases, resulting in subtle faults going…