2026
Teach2Eval: An Interaction-Driven LLMs Evaluation Method via Teaching Effectiveness
ICLR 2026poster
Recent progress in large language models (LLMs) has outpaced the development of effective evaluation methods. Evaluating LLMs with static, task-specific benchmarks is increasingly fragile due to contamination and saturation, and it fails to capture interactive reasoning. We introduce Teach2Eval, whi…