2026
FlashOptim: Memory Efficient Optimizers for Large-Scale Training
ICML 2026spotlight
Standard mixed-precision training of neural networks requires many bytes of accelerator memory for each model parameter. These bytes reflect not just the parameter itself, but also its gradient and one or more optimizer state variables. With each of these values typically requiring 4 bytes, training…