MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
The combination of Spiking Neural Networks (SNNs) with Vision Transformer architectures has attracted significant attention due to the great potential for energy-efficient and high-performance computing paradigms. However, a substantial performance gap still exists between SNN-based and ANN-based tr