2025
CoTD-PO: Chain-of-Thought Distillation with Preference Optimization
EMNLP 2025
Chain-of-Thought (CoT) distillation has emerged as a promising paradigm to enhance the reasoning ability of small language models by imitating the reasoning and outputs of larger teacher models. However, existing approaches suffer from a critical limitation: a distribution mismatch between teacher-g