2026
DeToNATION: Decoupled Torch Network-Aware Training on Interlinked Online Nodes
AAAI 2026technical
Training large neural network models requires extensive computational resources, often distributed across several nodes and accelerators. Recent findings suggest that it may be sufficient to only exchange the fast moving components of the gradients, while accumulating momentum locally (Decoupled Mom