Enhancing Small Model Performance in Educational Classification Tasks through Knowledge Distillation
Haoxin Xu, Changyong Qi, Bingqian Jiang, Tong Liu, Longwei Zheng, Xiaoqing Gu
Abstract
As the demand for precision, efficiency, and low-cost solutions in educational classification tasks continues to grow, enhancing model performance has become a critical focus of research. While large language models excel in these tasks, their high cost and resource requirements limit widespread application. This study proposes a Knowledge-Enhanced Distillation (KED) method, utilizing ChatGPT-4, ChatGPT-4o, and Llama3 as teacher models, and three different sizes of BERT models as student models. The method was validated across three real-world educational datasets. The results demonstrate that the KED method significantly improves the accuracy and F1 scores of small models in educational text classification tasks, while also substantially reducing computational costs and resource consumption. Notably, the KED method shows exceptional performance in scenarios involving few-shot learning and class imbalance. The innovation of this study lies in applying the KED method to educational classification tasks, filling a gap in current research and highlighting its significant potential for practical application in educational contexts.
BibTeX
@inproceedings{icassp2025_enhancingsmallmo,
title = {Enhancing Small Model Performance in Educational Classification Tasks through Knowledge Distillation},
author = {Haoxin Xu and Changyong Qi and Bingqian Jiang and Tong Liu and Longwei Zheng and Xiaoqing Gu},
booktitle = {ICASSP 2025},
year = {2025}
}