2026
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
ICLR 2026poster
The evaluation of code-generating Large Language Models (LLMs) is fundamentally constrained by two intertwined challenges: a reliance on static, easily contaminated problem sources and the use of superficial, low-rigor testing. This paper introduces a new benchmark construction philosophy, Dual Scal…