ICASSP 2025accepted0 citations

KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation

Shihao Wang, Zhengxing Huang, Enguang Zuo, Alimjan Aysa, Kurban Ubul

Abstract

Vision Transformer, with the distinctive architecture and self-attention mechanisms, had profoundly influenced the field of computer vision, establishing Transformer-based models as benchmarks for semantic segmentation. In this study, we propose a pioneering hybrid model that fuses Kolmogorov-Arnold convolutions with ViT architecture to tackle the intrinsic challenges of semantic segmentation. By leveraging the unique attributes of Kolmogorov-Arnold convolutions, our approach introduces a convolutional attention mechanism within the Vision Transformer framework, effectively alleviating the quadratic complexity associated with self-attention. Furthermore, we integrate large-kernel convolutions and an upsampling module into the decoder, which is designed to enhance feature resolution, capture fine details, and maintain robust performance in complex scenarios for dense prediction tasks. Comprehensive experiments conducted on the ADE20K, Cityscapes, and COCO-Stuff datasets reveal that our method achieves mean Intersection over Union (mIoU) scores of 55.52%, 83.6%, and 51.8%, respectively.

BibTeX
@inproceedings{icassp2025_kcgaformerwhenla,
  title = {KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation},
  author = {Shihao Wang and Zhengxing Huang and Enguang Zuo and Alimjan Aysa and Kurban Ubul},
  booktitle = {ICASSP 2025},
  year = {2025}
}
KCGAFormer: When Large-Kernel ConvFormer Meets KAN in Semantic Segmentation · ICASSP 2025