2026
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Kaijian Zou, Feiyang Xiong, Yunxiang Zhang, Xinliang Frederick Zhang, Yueqi Ren, Shitanshu Bhushan +3
ICML 2026poster
Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their complexity and ease of verification. Yet, current coding benchmarks face limitations such as lack of exceptionally challenging problems, insufficient test case cove…