← Search

Dmitry Yarotsky

13 accepted papers

2026

Corner Gradient Descent

ICLR 2026poster

We consider SGD-type optimization on infinite-dimensional quadratic problems with power law spectral conditions. It is well-known that on such problems deterministic GD has loss convergence rates $L_t=O(t^{-\zeta})$, which can be improved to $L_t=O(t^{-2\zeta})$ by using Heavy Ball with a non-statio…

Cited by 0SourcecodeScholar
2026

Gradient Flow Through Diagram Expansions: Learning Regimes and Explicit Solutions

ICML 2026spotlight

We develop a general mathematical framework to analyze scaling regimes and derive explicit analytic solutions for gradient flow (GF) in large learning problems. Our key innovation is a formal power series expansion of the loss evolution, with coefficients encoded by diagrams akin to Feynman diagrams…

Cited by 0SourceScholar
2023

A view of mini-batch SGD via generating functions: conditions of convergence, phase transitions, benefit from negative momenta.

ICLR 2023poster

Mini-batch SGD with momentum is a fundamental algorithm for learning large predictive models. In this paper we develop a new analytic framework to analyze noise-averaged properties of mini-batch SGD for linear models at constant learning rates, momenta and sizes of batches. Our key idea is to consid…

2023

Structure of universal formulas

NeurIPS 2023poster

By universal formulas we understand parameterized analytic expressions that have a fixed complexity, but nevertheless can approximate any continuous function on a compact set. There exist various examples of such formulas, including some in the form of neural networks. In this paper we analyze the e…

Cited by 1SourcePDFScholar
2021

Explicit loss asymptotics in the gradient descent training of neural networks

NeurIPS 2021poster

Current theoretical results on optimization trajectories of neural networks trained by gradient descent typically have the form of rigorous but potentially loose bounds on the loss values. In the present work we take a different approach and show that the learning trajectory of a wide network in a l…

Cited by 17SourcePDFScholar
2020

Theoretical Performance Bound of Uplink Channel Estimation Accuracy in Massive MIMO

ICASSP 2020accepted

In this paper, we present a new performance bound for uplink channel estimation (CE) accuracy in the Massive Multiple Input Multiple Output (MIMO) system. The proposed approach is based on noise power prediction after the CE unit. Our method outperforms the accuracy of a well-known Cramer-Rao lower…

Cited by 0SourceScholar