ICASSP 2025accepted0 citations

Two-stream Semantic Alignment Networks for Multi-label Image Classification

Wenlan Kuang, Zhixin Li

Abstract

Recent research ideas mainly focus on solving the semantic consistency of visual features and label features. However, since images contain complex scene content, the features captured by visual feature extraction networks based on grid or sequence representation may introduce redundant information or lack continuity when identifying irregular objects. In order to fully mine the visual information of complex objects in images and enhance the inter-modal interaction of images and labels, we introduce a flexible graph structure to explore the internal information of objects and design a Two-stream Semantic Alignment Networks (TsSAN) for multi-label image classification. To enhance the context awareness and semantic association of different patch regions, we propose a semantic-augmented interaction module that combines two kinds of visual semantic information with label embeddings for interactive learning. Finally, we refine the dependence between local intrinsic information and overall semantics by redefining semantic queries through semantically enhanced visual spatial features and graph aggregation features. Experiments on three public datasets demonstrate the effectiveness of our TsSAN.

BibTeX
@inproceedings{icassp2025_twostreamsemanti,
  title = {Two-stream Semantic Alignment Networks for Multi-label Image Classification},
  author = {Wenlan Kuang and Zhixin Li},
  booktitle = {ICASSP 2025},
  year = {2025}
}