2025
DOMAINEVAL: An Auto-Constructed Benchmark for Multi-Domain Code Generation
AAAI 2025technical
Code benchmarks such as HumanEval are widely adopted to evaluate the capabilities of Large Language Models (LLMs), providing insights into their strengths and weaknesses. However, current benchmarks primarily exercise LLMs' capability on common coding tasks (e.g., bubble sort, greatest common diviso…