← Search

Deanna Needell

11 accepted papers

2026

VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks

ICML 2026poster

We introduce vector diffusion wavelets (VDWs), a novel family of wavelets inspired by the vector diffusion maps algorithm that was introduced to analyze data lying in the tangent bundle of a Riemannian manifold. We show that these wavelets may be effectively incorporated into a family of geometric g…

Cited by 0SourceScholar
2025

Are Greedy Task Orderings Better Than Random in Continual Linear Regression?

NeurIPS 2025poster

We analyze task orderings in continual learning for linear regression, assuming joint realizability of training data. We focus on orderings that greedily maximize dissimilarity between consecutive tasks, a concept briefly explored in prior work but still surrounded by open questions. Using tools fro…

Cited by 0SourceScholar
2024

Benign overfitting in leaky ReLU networks with moderate input dimension

NeurIPS 2024spotlight

The problem of benign overfitting asks whether it is possible for a model to perfectly fit noisy training data and still generalize well. We study benign overfitting in two-layer leaky ReLU networks trained with the hinge loss on a binary classification task. We consider input data which can be deco…

Cited by 4SourcePDFScholar
2024

Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms

ICML 2024poster

We analyze inexact Riemannian gradient descent (RGD) where Riemannian gradients and retractions are inexactly (and cheaply) computed. Our focus is on understanding when inexact RGD converges and what is the complexity in the general nonconvex and constrained setting. We answer these questions in a g…

Cited by 0SourcePDFScholar
2023

Nearly Optimal Bounds for Cyclic Forgetting

NeurIPS 2023poster

We provide theoretical bounds on the forgetting quantity in the continual learning setting for linear tasks, where each round of learning corresponds to projecting onto a linear subspace. For a cyclic task ordering on $T$ tasks repeated $m$ times each, we prove the best known upper bound of $O(T^2/m…

Cited by 6SourcePDFScholar
2023

SP2 : A Second Order Stochastic Polyak Method

ICLR 2023poster

Recently the SP (Stochastic Polyak step size) method has emerged as a competitive adaptive method for setting the step sizes of SGD. SP can be interpreted as a method specialized to interpolated models, since it solves the interpolation equations. SP solves these equation by using local linearizati…

Cited by 13SourcePDFScholar
2023

Training shallow ReLU networks on noisy data using hinge loss: when do we overfit and is it benign?

NeurIPS 2023spotlight

We study benign overfitting in two-layer ReLU networks trained using gradient descent and hinge loss on noisy data for binary classification. In particular, we consider linearly separable data for which a relatively small proportion of labels are corrupted or flipped. We identify conditions on the m…

Cited by 9SourcePDFScholar