2025
Autonomous Evaluation of LLMs for Truth Maintenance and Reasoning Tasks
ICLR 2025poster
This paper presents AutoEval, a novel benchmark for scaling Large Language Model (LLM) assessment in formal tasks with clear notions of correctness, such as truth maintenance in translation and logical reasoning. AutoEval is the first benchmarking paradigm that offers several key advantages necessar…