2025
ToDi: Token-wise Distillation via Fine-Grained Divergence Control
EMNLP 2025
Large language models (LLMs) offer impressive performance but are impractical for resource-constrained deployment due to high latency and energy consumption. Knowledge distillation (KD) addresses this by transferring knowledge from a large teacher to a smaller student model. However, conventional KD