← Search

Zesen Wang

2 accepted papers

2025

From Promise to Practice: Realizing High-performance Decentralized Training

ICLR 2025poster

Decentralized training of deep neural networks has attracted significant attention for its theoretically superior scalability compared to synchronous data-parallel methods like All-Reduce. However, realizing this potential in multi-node training is challenging due to the complex design space that in…

2023

Bringing regularized optimal transport to lightspeed: a splitting method adapted for GPUs

NeurIPS 2023poster

We present an efficient algorithm for regularized optimal transport. In contrast to previous methods, we use the Douglas-Rachford splitting technique to develop an efficient solver that can handle a broad class of regularizers. The algorithm has strong global convergence guarantees, low per-iteratio…

Cited by 4SourcePDFScholar