2025
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
NeurIPS 2025poster
With the significant progress of large reasoning models in complex coding and reasoning tasks, existing benchmarks, like LiveCodeBench and CodeElo, are insufficient to evaluate the coding capabilities of large language models (LLMs) in real competition environments. Moreover, current evaluation met…