2026
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
ICLR 2026poster
We study the implicit bias of sharpness-aware minimization (SAM) when training $L$-layer linear diagonal networks on linearly separable binary classification. For linear models ($L=1$), both $\ell_\infty$- and $\ell_2$-SAM recover the $\ell_2$ max-margin classifier, matching gradient descent (GD). H…