← Search

Thinh On

1 accepted papers

2026

Knowledge Distillation for Large Language Models through Residual Learning

ICLR 2026poster

Knowledge distillation has become a crucial technique to transfer the capacities of large language models (LLMs) to smaller, more efficient models for practical deployment. While recent work exploits rich information from intermediate states of the teacher model for more effective knowledge transfer…

Cited by 0SourceScholar