2026
Beyond the Mean: Fisher-Orthogonal Projection for Natural Gradient Descent in Large Batch Training
AAAI 2026technical
Modern GPUs are equipped with large amounts of high-bandwidth memory, enabling them to support mini-batch sizes of up to tens of thousands of training samples. However, most existing optimizers struggle to perform effectively at such a large batch size. As batch size increases, gradient noise decre