ICASSP 2025accepted0 citations

Multi-Task Joint 3D Swin Transformer Learning for Segmentation and Classification of Hyperspectral Medicine Images

Dong Zhang, Meijun Sun

Abstract

Hyperspectral images had made many applications in the medical field with their rich spectral information. However, there were currently problems with feature extraction based on hyperspectral images, especially in extracting contextual feature information from spectral bands, and a single convolutional kernel may restrict the receptive field and inadequately capture the sequential properties of the data. Meanwhile, due to the large data volume of hyperspectral images, current medical hyperspectral images focus more on individual segmentation or classification. This paper proposed a 3D swin transformer with multi-task joint learning framework, to simultaneously learn multiple tasks for hyperspectral tongue images. Based on the 3D swin transformer model, the framework regards cross-band context feature learning of hyperspectral images as a sequence-to-sequence prediction process. The 3D transformer encoder used as the basic framework for shared feature extraction, and set up corresponding decoders for prediction according to different visual tasks such as segmentation and classification. We conducted simultaneous segmentation and classification tasks of the tongue coating region on the hyperspectral image of tongue images. The results showed that the proposed model had good results in single tasks of segmentation and classification, and performed better than other multi-task convolutional neural network models.

BibTeX
@inproceedings{icassp2025_multitaskjoint3d,
  title = {Multi-Task Joint 3D Swin Transformer Learning for Segmentation and Classification of Hyperspectral Medicine Images},
  author = {Dong Zhang and Meijun Sun},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Multi-Task Joint 3D Swin Transformer Learning for Segmentation and Classification of Hyperspectral Medicine Images · ICASSP 2025