← Search

Egor Petrov

3 accepted papers

2026

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

ICML 2026poster

Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations leave GPUs idle during pipeline bubbles, wasting computational resources. Asynchronous Pipeline Parallelism approaches effectively eliminate these bubbles, maximizing throughput at t…

Cited by 0SourceScholar
2026

Sign-SGD via Parameter-Free Optimization

ICLR 2026poster

Large language models have achieved major advances across domains, yet training them remains extremely resource-intensive. We revisit Sign-SGD, which serves both as a memory-efficient optimizer for single-node training and as a gradient compression mechanism for distributed learning. This paper addr…

Cited by 0SourceScholar
2025

When Extragradient Meets PAGE: Bridging Two Giants to Boost Variational Inequalities

UAI 2025

Variational inequalities (VIs) have emerged as a universal framework for solving a wide range of problems. A broad spectrum of applications include optimization, equilibrium analysis, reinforcement learning, and the rapidly evolving field of generative adversarial networks (GANs). Stochastic methods

Cited by 0SourcePDFScholar