← Search

Yijun Dong

7 accepted papers

2026

A Task-centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula

ICML 2026poster

Iterative self-improvement fine-tunes an autoregressive large language model (LLM) on reward-verified outputs generated by the LLM itself. In contrast to the empirical success of self-improvement, the theoretical foundation of this generative, iterative procedure in a practical, finite-sample settin…

Cited by 0SourceScholar
2025

Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic Dimension

ICML 2025poster

Weak-to-strong (W2S) generalization is a type of finetuning (FT) where a strong (large) student model is trained on pseudo-labels generated by a weak teacher. Surprisingly, W2S FT often outperforms the weak teacher. We seek to understand this phenomenon through the observation that FT often occurs i…

Cited by 0SourcePDFScholar
2024

Sketchy Moment Matching: Toward Fast and Provable Data Selection for Finetuning

NeurIPS 2024poster

We revisit data selection in a modern context of finetuning from a fundamental perspective. Extending the classical wisdom of variance minimization in low dimensions to high-dimensional finetuning, our generalization analysis unveils the importance of additionally reducing bias induced by low-rank a…

2023

Adaptively Weighted Data Augmentation Consistency Regularization for Robust Optimization under Concept Shift

ICML 2023poster

Concept shift is a prevailing problem in natural tasks like medical image segmentation where samples usually come from different subpopulations with variant correlations between features and labels. One common type of concept shift in medical image segmentation is the "information imbalance" between…

Cited by 3SourcePDFScholar
2023

Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering

NeurIPS 2023poster

Despite the empirical success and practical significance of (relational) knowledge distillation that matches (the relations of) features between teacher and student models, the corresponding theoretical interpretations remain limited for various knowledge distillation paradigms. In this work, we tak…

Cited by 6SourcePDFScholar
2023

Sample Efficiency of Data Augmentation Consistency Regularization

AISTATS 2023poster

Data augmentation is popular in the training of large neural networks; however, currently, theoretical understanding of the discrepancy between different algorithmic choices of leveraging augmented data remains limited. In this paper, we take a step in this direction – we first present a simple and…

Cited by 25SourcePDFScholar