2025
Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic Activations
NeurIPS 2025poster
The theory of training deep networks has become a central question of modern machine learning and has inspired many practical advancements. In particular, the gradient descent (GD) optimization algorithm has been extensively studied in recent years. A key assumption about GD has appeared in several…