2025
DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language Models
ICML 2025poster
Although knowledge distillation (KD) is an effective approach to improve the performance of a smaller LLM (i.e., the student model) by transferring knowledge from a large LLM (i.e., the teacher model), it still suffers from high training cost. Existing LLM distillation methods ignore the difficulty…