← Search

Bingqing Song

3 accepted papers

2024

Unraveling the Gradient Descent Dynamics of Transformers

NeurIPS 2024poster

While the Transformer architecture has achieved remarkable success across various domains, a thorough theoretical foundation explaining its optimization dynamics is yet to be fully developed. In this study, we aim to bridge this understanding gap by answering the following two core questions: (1) Wh…

Cited by 1SourcePDFScholar
2023

FedAvg Converges to Zero Training Loss Linearly for Overparameterized Multi-Layer Neural Networks

ICML 2023poster

Federated Learning (FL) is a distributed learning paradigm that allows multiple clients to learn a joint model by utilizing privately held data at each client. Significant research efforts have been devoted to develop advanced algorithms that deal with the situation where the data at individual clie…

Cited by 8SourcePDFScholar
2022

Distributed Optimization for Overparameterized Problems: Achieving Optimal Dimension Independent Communication Complexity

NeurIPS 2022accept

Decentralized optimization are playing an important role in applications such as training large machine learning models, among others. Despite its superior practical performance, there has been some lack of fundamental understanding about its theoretical properties. In this work, we address the foll…

Cited by 5SourcePDFScholar