2025
Humanity’s Last Code Exam: Can Advanced LLMs Conquer Human’s Hardest Code Competition?
EMNLP 2025
Code generation is a core capability of large language models (LLMs), yet mainstream benchmarks (e.g., APPs and LiveCodeBench) contain questions with medium-level difficulty and pose no challenge to advanced LLMs. To better reflected the advanced reasoning and code generation ability, We introduce H