← Search

Ji Zeng

2 accepted papers

2026

QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

ICML 2026poster

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate …

Cited by 1SourceScholar
2026

daVinci-Dev: Agent-native Mid-training for Software Engineering

ICML 2026oral

Recently, the frontier of Large Language Model (LLM) capabilities has shifted from single-turn code generation to agentic software engineering—a paradigm where models autonomously navigate, edit, and test complex repositories. While post-training methods have become the de facto approach for code ag…

Cited by 0SourceScholar