← Search

Todor Antić

1 accepted papers

2026

QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

ICML 2026poster

As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate …

Cited by 0SourceScholar