← Search

Maxim Kodryan

5 accepted papers

2024

Where Do Large Learning Rates Lead Us?

NeurIPS 2024poster

It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is r…

2023

MARS: Masked Automatic Ranks Selection in Tensor Decompositions

AISTATS 2023poster

Tensor decomposition methods have proven effective in various applications, including compression and acceleration of neural networks. At the same time, the problem of determining optimal decomposition ranks, which present the crucial parameter controlling the compressionaccuracy trade-off, is still…

2022

Training Scale-Invariant Neural Networks on the Sphere Can Happen in Three Regimes

NeurIPS 2022accept

A fundamental property of deep learning normalization techniques, such as batch normalization, is making the pre-normalization parameters scale invariant. The intrinsic domain of such parameters is the unit sphere, and therefore their gradient optimization dynamics can be represented via spherical o…

2021

On the Periodic Behavior of Neural Network Training with Batch Normalization and Weight Decay

NeurIPS 2021poster

Training neural networks with batch normalization and weight decay has become a common practice in recent years. In this work, we show that their combined use may result in a surprising periodic behavior of optimization dynamics: the training process regularly exhibits destabilizations that, however…

2020

On Power Laws in Deep Ensembles

NeurIPS 2020spotlight

Ensembles of deep neural networks are known to achieve state-of-the-art performance in uncertainty estimation and lead to accuracy improvement. In this work, we focus on a classification problem and investigate the behavior of both non-calibrated and calibrated negative log-likelihood (CNLL) of a de…