2025
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
NeurIPS 2025poster
The scaling law for large language models (LLMs) depicts that the path towards machine intelligence necessitates training at large scale. Thus, companies continuously build large-scale GPU clusters, and launch training jobs that span over thousands of computing nodes. However, LLM pre-training prese…