2024
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
EMNLP 2024finding
Driven by the surge in code generation using large language models (LLMs), numerous benchmarks have emerged to evaluate these LLMs capabilities. We conducted a large-scale human evaluation of *HumanEval* and *MBPP*, two popular benchmarks for Python code generation, analyzing their diversity and dif…