ICASSP 2025accepted0 citations

Text-Guided Few-Shot Semantic Segmentation with Training-Free Multimodal Feature Matching

Guillaume Buthmann, Tomoya Sakai, Haoxiang Qiu, Takayuki Katsuki, Daiki Kimura

Abstract

This paper addresses few-shot semantic segmentation (FSS) guided by text, where we classify unseen novel classes using image and text references as in-context examples, without the need for training. We enhance the quality and stability of the segmentation masks generated by FSS by combining the capability of open-vocabulary zero-shot semantic segmentation (ZSS) based on foundation models for image and text. We propose a training-free approach using multimodal feature matching that performs segmentation by identifying regions in a target image that match the features from both the image and text references. Experimental results demonstrate that the proposed method outperforms state-of-the-art FSS and ZSS methods.

BibTeX
@inproceedings{icassp2025_textguidedfewsho,
  title = {Text-Guided Few-Shot Semantic Segmentation with Training-Free Multimodal Feature Matching},
  author = {Guillaume Buthmann and Tomoya Sakai and Haoxiang Qiu and Takayuki Katsuki and Daiki Kimura},
  booktitle = {ICASSP 2025},
  year = {2025}
}