"Idling Neurons, Appropriately Lenient Workload During Fine-tuning Leads to Better Generalization"
"Pre-training on large-scale datasets has become a fundamental method for training deep neural networks. Pre-training provides a better set of parameters than random initialization, which reduces the training cost of deep neural networks on the target task. In addition, pre-training also provides a…