2025
BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
ICLR 2025oral
Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks range from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human develop…