← Search

Jiangcheng Song

3 accepted papers

2026

SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs

CVPR 2026

Knowledge distillation (KD) is a standard route to compress Large Language Models (LLMs) into compact students, yet most pipelines uniformly apply token-wise loss regardless of teacher confidence. This indiscriminate supervision amplifies noisy, high-entropy signals and is especially harmful under l

Cited by 0SourcecodeScholar
2026

“The Whole Is Greater than the Sum of Its Parts”: A Compatibility-Aware Multi-Teacher CoT Distillation Framework

IJCAI 2026

Chain-of-Thought (CoT) reasoning empowers Large Language Models (LLMs) with remarkable capabilities but typically requires prohibitive parameter scales. CoT distillation has emerged as a promising paradigm to transfer reasoning prowess into compact Student Models (SLMs), but existing approaches ofte

Cited by 0Scholar
2025

DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation Trainer

NeurIPS 2025poster

Recent advances in knowledge distillation have emphasized the importance of decoupling different knowledge components. While existing methods utilize momentum mechanisms to separate task-oriented and distillation gradients, they overlook the inherent conflict between target-class and non-target-clas…

Cited by 1SourcecodeScholar