2024
Flag Aggregator: Scalable Distributed Training under Failures and Augmented Losses using Convex Optimization
ICLR 2024poster
Modern ML applications increasingly rely on complex deep learning models and large datasets. There has been an exponential growth in the amount of computation needed to train the largest models. Therefore, to scale computation and data, these models are inevitably trained in a distributed manner in…