ICASSP 2025accepted0 citations

SlotFusion: Object-Centric Audiovisual Feature Fusion with Slot Attention for Remote Sensing Scene Recognition

Fangzhou Han, Tianyi Yu, Lamei Zhang, Lingyu Si, Yiqi Zhang

Abstract

Despite significant advancements in remote sensing multimodal learning, particularly in image-image feature fusion, the exploration of audio-image feature fusion remains insufficient. Given the complexity and redundancy of ground objects in remote sensing images, accurately aligning audio features with image features during the fusion process is a critical challenge. In this paper, we introduce an object-centric feature fusion method named SlotFusion. By employing a slot attention-based feature decoupling module and a slot-based audiovisual feature fusion module, we transform modality features with complex semantic information into a set of slot features corresponding to object units and use gated activation units to adaptively implement object-centric feature fusion. Experiments on the Audio Visual Aerial Scene Recognition dataset (ADVANCE) demonstrate that the proposed SlotFusion significantly improves remote sensing scene recognition performance, with a 7.04% increase in overall accuracy compared to previous methods, achieving state-of-the-art results.

BibTeX
@inproceedings{icassp2025_slotfusionobject,
  title = {SlotFusion: Object-Centric Audiovisual Feature Fusion with Slot Attention for Remote Sensing Scene Recognition},
  author = {Fangzhou Han and Tianyi Yu and Lamei Zhang and Lingyu Si and Yiqi Zhang},
  booktitle = {ICASSP 2025},
  year = {2025}
}