← Search

Rachel Ward

11 accepted papers

2024

Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural Networks

NeurIPS 2024poster

We study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solut…

Cited by 1SourcePDFScholar
2023

Adaptively Weighted Data Augmentation Consistency Regularization for Robust Optimization under Concept Shift

ICML 2023poster

Concept shift is a prevailing problem in natural tasks like medical image segmentation where samples usually come from different subpopulations with variant correlations between features and labels. One common type of concept shift in medical image segmentation is the "information imbalance" between…

Cited by 1SourcePDFScholar
2023

Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns Clustering

NeurIPS 2023poster

Despite the empirical success and practical significance of (relational) knowledge distillation that matches (the relations of) features between teacher and student models, the corresponding theoretical interpretations remain limited for various knowledge distillation paradigms. In this work, we tak…

Cited by 6SourcePDFScholar
2023

Nearly Optimal Bounds for Cyclic Forgetting

NeurIPS 2023poster

We provide theoretical bounds on the forgetting quantity in the continual learning setting for linear tasks, where each round of learning corresponds to projecting onto a linear subspace. For a cyclic task ordering on $T$ tasks repeated $m$ times each, we prove the best known upper bound of $O(T^2/m…

Cited by 6SourcePDFScholar
2023

Sample Efficiency of Data Augmentation Consistency Regularization

AISTATS 2023poster

Data augmentation is popular in the training of large neural networks; however, currently, theoretical understanding of the discrepancy between different algorithmic choices of leveraging augmented data remains limited. In this paper, we take a step in this direction – we first present a simple and…

Cited by 25SourcePDFScholar
2022

AdaLoss: A Computationally-Efficient and Provably Convergent Adaptive Gradient Method

AAAI 2022technical

We propose a computationally-friendly adaptive learning rate schedule, ``AdaLoss", which directly uses the information of the loss function to adjust the stepsize in gradient descent methods. We prove that this schedule enjoys linear convergence in linear regression. Moreover, we extend the to the n…

2020

Implicit Regularization and Convergence for Weight Normalization

NeurIPS 2020poster

Normalization methods such as batch, weight, instance, and layer normalization are commonly used in modern machine learning. Here, we study the weight normalization (WN) method \cite{salimans2016weight} and a variant called reparametrized projected gradient descent (rPGD) for overparametrized least…

Cited by 26SourcePDFScholar