← Search

Catherine Li

1 accepted papers

2026

Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting

ICML 2026poster

Standard optimizer choices for pre-training are designed to minimize pre-training loss. Yet pre-trained models are routinely subjected to further transformations—such as fine-tuning to acquire new capabilities or quantization for efficiency. In this work, we evaluate optimizer choices across model s…

Cited by 0SourceScholar