← Search

Thomas Pethick

11 accepted papers

2025

Efficient Interpolation between Extragradient and Proximal Methods for Weak MVIs

ICLR 2025poster

We study nonmonotone games satisfying the weak Minty variational inequality (MVI) with parameter $\rho \in (-\tfrac{1}{L}, \infty)$, where $L$ is the Lipschitz constant of the gradient operator. An error corrected version of the inexact proximal point algorithm is proposed, with which we establish t…

Cited by 0SourcePDFScholar
2025

Generalized Gradient Norm Clipping & Non-Euclidean $(L_0,L_1)$-Smoothness

NeurIPS 2025oral

This work introduces a hybrid non-Euclidean optimization method which generalizes gradient norm clipping by combining steepest descent and conditional gradient approaches. The method achieves the best of both worlds by establishing a descent property under a generalized notion of ($L_0$,$L_1$)-smoot…

Cited by 0SourcecodeScholar
2025

Training Deep Learning Models with Norm-Constrained LMOs

ICML 2025spotlight

In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball. We propose a new stochastic family of algorithms that uses the LMO to adapt to the geometry of the problem and, perhaps surprisingly, show that they can be applied to unconstrained proble…

2024

Improving SAM Requires Rethinking its Optimization Formulation

ICML 2024poster

This paper rethinks Sharpness-Aware Minimization (SAM), which is originally formulated as a zero-sum game where the weights of a network and a bounded perturbation try to minimize/maximize, respectively, the same differentiable loss. To fundamentally improve this design, we argue that SAM should ins…

2023

Finding Actual Descent Directions for Adversarial Training

ICLR 2023poster

Adversarial Training using a strong first-order adversary (PGD) is the gold standard for training Deep Neural Networks that are robust to adversarial examples. We show that, contrary to the general understanding of the method, the gradient at an optimal adversarial example may increase, rather than…

Cited by 0SourcePDFScholar
2023

Solving stochastic weak Minty variational inequalities without increasing batch size

ICLR 2023poster

This paper introduces a family of stochastic extragradient-type algorithms for a class of nonconvex-nonconcave problems characterized by the weak Minty variational inequality (MVI). Unlike existing results on extragradient methods in the monotone setting, employing diminishing stepsizes is no longer…

2023

Stable Nonconvex-Nonconcave Training via Linear Interpolation

NeurIPS 2023spotlight

This paper presents a theoretical analysis of linear interpolation as a principled method for stabilizing (large-scale) neural network training. We argue that instabilities in the optimization process are often caused by the nonmonotonicity of the loss landscape and show how linear interpolation can…

2022

Escaping limit cycles: Global convergence for constrained nonconvex-nonconcave minimax problems

ICLR 2022spotlight

This paper introduces a new extragradient-type algorithm for a class of nonconvex-nonconcave minimax problems. It is well-known that finding a local solution for general minimax problems is computationally intractable. This observation has recently motivated the study of structures sufficient for co…

2021

Sifting through the noise: Universal first-order methods for stochastic variational inequalities

NeurIPS 2021poster

We examine a flexible algorithmic framework for solving monotone variational inequalities in the presence of randomness and uncertainty. The proposed template encompasses a wide range of popular first-order methods, including dual averaging, dual extrapolation and optimistic gradient algorithms – bo…

Cited by 13SourcePDFScholar
2021

Subquadratic Overparameterization for Shallow Neural Networks

NeurIPS 2021poster

Overparameterization refers to the important phenomenon where the width of a neural network is chosen such that learning algorithms can provably attain zero loss in nonconvex training. The existing theory establishes such global convergence using various initialization strategies, training modificat…

Cited by 36SourcePDFScholar