Memory-Efficient LLMs Training with Dynamic Sparsity: From Stability to Practical Scaling
Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST suffers from optimization instability, manifested as loss spikes following topology updates. In this work…