← Search

Dachao Lin

5 accepted papers

2023

Stochastic Distributed Optimization under Average Second-order Similarity: Algorithms and Analysis

NeurIPS 2023poster

We study finite-sum distributed optimization problems involving a master node and $n-1$ local nodes under the popular $\delta$-similarity and $\mu$-strong convexity conditions. We propose two new algorithms, SVRS and AccSVRS, motivated by previous works. The non-accelerated SVRS method combines the…

Cited by 13SourcePDFScholar
2021

Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data

NeurIPS 2021poster

In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate for (deep) linear networks with spherically symmetric data distribution, which ca…

Cited by 4SourcePDFScholar
2021

Greedy and Random Quasi-Newton Methods with Faster Explicit Superlinear Convergence

NeurIPS 2021poster

In this paper, we follow Rodomanov and Nesterov’s work to study quasi-Newton methods. We focus on the common SR1 and BFGS quasi-Newton methods to establish better explicit (local) superlinear convergence rates. First, based on the greedy quasi-Newton update which greedily selects the direction to ma…

Cited by 19SourcePDFScholar
2019

Toward Understanding the Importance of Noise in Training Neural Networks

ICML 2019oral

Numerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of deep neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem through training a simple two-layer convolutional neural net…

Cited by 106SourcePDFScholar