← Search

Xiaoge Deng

5 accepted papers

2025

Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks

ICASSP 2025accepted

Sharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated with the regularization to characterize the sharpness of the model. Despite its success, research on adaptive regularizat…

Cited by 0SourceScholar
2024

Exploring the Inefficiency of Heavy Ball as Momentum Parameter Approaches 1

IJCAI 2024poster

The heavy ball momentum method is a commonly used technique for accelerating training processes in the machine learning community. However, empirical evidence suggests that the convergence of stochastic gradient descent (SGD) with heavy ball may slow down when the momentum hyperparameter approaches…

Cited by 0SourcePDFScholar
2024

Stability and Generalization of Asynchronous SGD: Sharper Bounds Beyond Lipschitz and Smoothness

NeurIPS 2024poster

Asynchronous stochastic gradient descent (ASGD) has evolved into an indispensable optimization algorithm for training modern large-scale distributed machine learning tasks. Therefore, it is imperative to explore the generalization performance of the ASGD algorithm. However, the existing results are…

Cited by 4SourcePDFScholar
2023

Stability-Based Generalization Analysis of the Asynchronous Decentralized SGD

AAAI 2023technical

The generalization ability often determines the success of machine learning algorithms in practice. Therefore, it is of great theoretical and practical importance to understand and bound the generalization error of machine learning algorithms. In this paper, we provide the first generalization resul…

Cited by 19SourcePDFScholar
2022

S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning

ICASSP 2022accepted

Distributed stochastic gradient descent (SGD) approach has been widely used in large-scale deep learning, and the gradient collective method is vital to ensure the training scalability of the distributed deep learning system. Collective communication such as AllReduce has been widely adopted for the…

Cited by 0SourceScholar