2025
Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension
ICML 2025poster
Weak-to-strong (W2S) generalization is a type of finetuning (FT) where a strong (large) student model is trained on pseudo-labels generated by a weak teacher. Surprisingly, W2S FT often outperforms the weak teacher. We seek to understand this phenomenon through the observation that FT often occurs i…