ICLR 2026poster0 citations

Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching Effectiveness

Yuhang Zhou, Xutian Chen, Yixin Cao, Yuchen Ni, Yu He, Siyu Tian, Xiang Liu, Yunwen Chen

Abstract

Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Evaluating LLMs with static, task-specific benchmarks is increasingly fragile due to contamination and saturation, and it fails to capture interactive reasoning. We introduce Teach2Eval, which reframes evaluation as teaching: a candidate model guides weaker students, and the students’ gains constitute the score. This interaction yields robustness to contamination and exposes orthogonal abilities with fine-grained metrics across Application, Judgment, Guidance, and Reflection. The framework scales automatically by exploiting natural error distributions from weak students, requiring neither bespoke rubrics nor human graders. Across 30 LLMs and 60 datasets, Teach2Eval achieves Spearman above 0.95 with human-preference leaderboards (e.g., Chatbot Arena/LiveBench), surpassing direct baselines, while offering actionable training signals (capability hierarchies, early overfitting) at low cost.

New Evaluation MethodMulti-dimensional EvaluationLarge Language ModelsData ContaminationTeach2Eval
BibTeX
@inproceedings{
zhou2026teacheval,
title={Teach2Eval: An Interaction-Driven {LLM}s Evaluation Method via Teaching Effectiveness},
author={Yuhang Zhou and Xutian Chen and Yixin Cao and Yuchen Ni and Yu He and Siyu Tian and Xiang Liu and Yunwen Chen and Guangnan Ye and Xipeng Qiu and Hongfeng Chai},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=HreYquZ5xs}
}