2026
From Solver to Tutor: Evaluating the Pedagogical Intelligence of LLMs with KMP-Bench
AAAI 2026technical
Large Language Models (LLMs) show significant potential in AI mathematical tutoring, yet current evaluations often rely on simplistic metrics or narrow pedagogical scenarios, failing to assess comprehensive, multi-turn teaching effectiveness. In this paper, we introduce KMP-Bench, a comprehensive K-