← Search

Ziqing Xu

4 accepted papers

2025

Convergence Rates for Gradient Descent on the Edge of Stability for Overparametrised Least Squares

NeurIPS 2025poster

Classical optimisation theory guarantees monotonic objective decrease for gradient descent (GD) when employed in a small step size, or "stable", regime. In contrast, gradient descent on neural networks is frequently performed in a large step size regime called the "edge of stability", in which the o…

Cited by 0SourceScholar
2025

Understanding the Learning Dynamics of LoRA: A Gradient Flow Perspective on Low-Rank Adaptation in Matrix Factorization

AISTATS 2025poster

Despite the empirical success of Low-Rank Adaptation (LoRA) in fine-tuning pre-trained models, there is little theoretical understanding of how first-order methods with carefully crafted initialization adapt models to new tasks. In this work, we take the first step towards bridging this gap by theor…

Cited by 0SourceScholar
2023

Linear Convergence of Gradient Descent For Finite Width Over-parametrized Linear Networks With General Initialization

AISTATS 2023poster

Recent theoretical analyses of the convergence of gradient descent (GD) to a global minimum for over-parametrized neural networks make strong assumptions on the step size (infinitesimal), the hidden-layer width (infinite), or the initialization (spectral, balanced). In this work, we relax these assu…

Cited by 8SourcePDFScholar
2021

Deep Transfer Tensor Decomposition with Orthogonal Constraint for Recommender Systems

AAAI 2021technical

Tensor decomposition is one of the most effective techniques for multi-criteria recommendations. However, it suffers from data sparsity when dealing with three-dimensional (3D) user-item-criterion ratings. To mitigate this issue, we consider effectively incorporating the side information and cross-d…

Cited by 52SourcePDFScholar