2026
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators
ICLR 2026poster
Large Language Models (LLMs) have shown impressive performance across diverse domains, with code generation emerging as a particularly prominent application. However, existing benchmarks designed to evaluate code generation exhibit several critical limitations. First, most rely on manual annotations…