2026
WBMM: Windowed Batch Matrix Multiplication for Efficient Large Receptive Field Convolution
ICML 2026spotlight
Large kernel depthwise convolutions achieve strong performance but suffer from significant degradation as kernel size grows due to irregular memory access from gather-based computation. While Large Kernel Acceleration (LKA) helps on small feature maps, it becomes \textbf{counterproductive on large f…