2026
Critique-Guided Distillation for Robust Reasoning via Refinement
ICML 2026poster
Supervised fine-tuning with expert demonstrations often produces models that imitate outputs without internalizing the reasoning processes needed for robust generalization. While critique-based approaches show promise, training models to generate critiques directly, such as Critique Fine-Tuning (CFT…