2016
Training Neural Networks Without Gradients: A Scalable ADMM Approach
ICML 2016poster
With the growing importance of large network models and enormous training datasets, GPUs have become increasingly necessary to train neural networks. This is largely because conventional optimization algorithms rely on stochastic gradient methods that don’t scale well to large numbers of cores in a…