← Search

Nazarii Tupitsa

6 accepted papers

2025

Methods for Convex $(L_0,L_1)$-Smooth Optimization: Clipping, Acceleration, and Adaptivity

ICLR 2025poster

Due to the non-smoothness of optimization problems in Machine Learning, generalized smoothness assumptions have been gaining a lot of attention in recent years. One of the most popular assumptions of this type is $(L_0,L_1)$-smoothness (Zhang et al., 2020). In this paper, we focus on the class of (s…

Cited by 17SourcePDFScholar
2024

Low-Resource Machine Translation through the Lens of Personalized Federated Learning

EMNLP 2024finding

We present a new approach called MeritOpt based on the Personalized Federated Learning algorithm MeritFed that can be applied to Natural Language Tasks with heterogeneous data. We evaluate it on the Low-Resource Machine Translation task, using the datasets of South East Asian and Finno-Ugric languag…

2024

Remove that Square Root: A New Efficient Scale-Invariant Version of AdaGrad

NeurIPS 2024poster

Adaptive methods are extremely popular in machine learning as they make learning rate tuning less expensive. This paper introduces a novel optimization algorithm named KATE, which presents a scale-invariant adaptation of the well-known AdaGrad algorithm. We prove the scale-invariance of KATE for the…

2023

Byzantine-Tolerant Methods for Distributed Variational Inequalities

NeurIPS 2023poster

Robustness to Byzantine attacks is a necessity for various distributed training scenarios. When the training reduces to the process of solving a minimization problem, Byzantine robustness is relatively well-understood. However, other problem formulations, such as min-max problems or, more generally,…

Cited by 0SourcePDFScholar
2021

On a Combination of Alternating Minimization and Nesterov’s Momentum

ICML 2021spotlight

Alternating minimization (AM) procedures are practically efficient in many applications for solving convex and non-convex optimization problems. On the other hand, Nesterov’s accelerated gradient is theoretically optimal first-order method for convex optimization. In this paper we combine AM and Nes…

2019

On the Complexity of Approximating Wasserstein Barycenters

ICML 2019oral

We study the complexity of approximating the Wasserstein barycenter of $m$ discrete measures, or histograms of size $n$, by contrasting two alternative approaches that use entropic regularization. The first approach is based on the Iterative Bregman Projections (IBP) algorithm for which our novel an…

Cited by 124SourcePDFScholar