2025
Automated Benchmark Generation for Repository-Level Coding Tasks
ICML 2025poster
Code Agent development is an extremely active research area, where a reliable performance metric is critical for tracking progress and guiding new developments. This demand is underscored by the meteoric rise in popularity of SWE-Bench -- a benchmark that challenges code agents to generate patches a…