2018
Gradients explode - Deep Networks are shallow - ResNet explained
ICLR 2018workshop
Whereas it is believed that techniques such as Adam, batch normalization and, more recently, SeLU nonlinearities ``solve'' the exploding gradient problem, we show that this is not the case and that in a range of popular MLP architectures, exploding gradients exist and that they limit the depth to wh…