2026
Knowledge Distillation for Large Language Models through Residual Learning
ICLR 2026poster
Knowledge distillation has become a crucial technique to transfer the capacities of large language models (LLMs) to smaller, more efficient models for practical deployment. While recent work exploits rich information from intermediate states of the teacher model for more effective knowledge transfer…