ICML 2026poster0 citations

Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo

Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Hadi Mohaghegh Dolatabadi, Shamane Siriwardhana, Gil Avraham, Violetta Shevchenko, Karol Pajak

Abstract

To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently, its outer synchronization remains bandwidth-heavy and brittle to stragglers and transient failures. We relax exact synchronization to approximate synchronization via mixing/gossip, which degrades gracefully under delays and communication failures. This allows us to factorize DiLoCo synchronization into a non-blocking mixing step that overlaps computation with no staleness, and a blocking mixing step that tightens worker agreement, yielding a tunable trade-off between compute utilization and optimization stability. On up to billion-parameter language models in low-bandwidth settings, our method substantially improves compute utilization while matching DiLoCo’s training progress, and is more robust to failures.

OptimizationRobustnessRetrieval
BibTeX
@inproceedings{
koneputugodage2026factored,
title={Factored Gossip DiLoCo: Reducing Blocking Communication within DiLoCo},
author={Chamin P Hewa Koneputugodage and Thalaiyasingam Ajanthan and Sameera Ramasinghe and Hadi Mohaghegh Dolatabadi and Shamane Siriwardhana and Gil Avraham and Violetta Shevchenko and Karol Pajak and James Snewin and Alexander Long},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=1MynM9hkFH}
}