2024
DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories
ACL 2024findings
How to evaluate the coding abilities of Large Language Models (LLMs) remains an open question. We find that existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of LLMs.To address the knowledge gap, we propose a new benchmark…