2025
HumanEval Pro and MBPP Pro: Evaluating Large Language Models on Self-invoking Code Generation Task
ACL 2025finding
In this paper, we present HumanEval Pro and MBPP Pro, a series of benchmarks to evaluate LLMs on self-invoking code generation task. This task involves providing LLMs with a base problem alongside a related, more complex problem. The models must solve the base problem and leverage its solution to ad…