2025
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
ICLR 2025poster
Large Language Models (LLMs) applied to code-related applications have emerged as a prominent field, attracting significant interest from academia and industry. However, as new and improved LLMs are developed, existing evaluation benchmarks (e.g., HumanEvla, MBPP) are no longer sufficient for assess…