← Search

Ildus Sadrtdinov

3 accepted papers

2026

Can Stationary Distributions of Scale-Invariant Neural Networks Be Described by the Thermodynamics of an Ideal Gas?

IJCAI 2026

Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights. Building on this perspective, we develop a thermodynamic framework to describe the stationary distributions of stochastic gradient descent (SGD) wit

Cited by 0Scholar
2024

Where Do Large Learning Rates Lead Us?

NeurIPS 2024poster

It is generally accepted that starting neural networks training with large learning rates (LRs) improves generalization. Following a line of research devoted to understanding this effect, we conduct an empirical study in a controlled setting focusing on two questions: 1) how large an initial LR is r…

2023

To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning

NeurIPS 2023poster

Transfer learning and ensembling are two popular techniques for improving the performance and robustness of neural networks. Due to the high cost of pre-training, ensembles of models fine-tuned from a single pre-trained checkpoint are often used in practice. Such models end up in the same basin of…