2025
Enhancing Knowledge Distillation of Large Language Models through Efficient Multi-Modal Distribution Alignment
COLING 2025main
Knowledge distillation (KD) is an effective model compression method that can transfer the internal capabilities of large language models (LLMs) to smaller ones. However, the multi-modal probability distribution predicted by teacher LLMs causes difficulties for student models to learn. In this paper…