ICASSP 2025accepted0 citations

Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation

Shuo Zhang, Jiaming Huang, Yan Wu, Tao Hu, Wenbing Tang, Jing Liu

Abstract

Open-vocabulary semantic segmentation (OVSS) is a challenging computer vision task that labels each pixel within an image based on text descriptions. Recent advancements in OVSS are largely attributed to the increased model capacity. However, these models often struggle with unfamiliar images or unseen text, as their visual language understanding is limited to training data. Text-to-image (T2I) diffusion models have demonstrated strong image generation with diverse open-vocabulary descriptions. It prompted us to explore whether the comprehensive priors in T2I diffusion models could enhance the zero-shot generalization of OVSS. In this study, we define OVSS as a denoising diffusion task from noisy to object mask and introduce Seg-diffusion, a novel method based on Stable Diffusion that utilizes its extensive visual and linguistic prior knowledge. Specifically, the object mask diffuses from ground-truth to a random distribution in latent space. The model learns to reverse this noisy process to reconstruct object mask to segment objectives using text embeddings with our proposed Content Attention Module (CAM). Extensive experiments on popular OVSS benchmarks show that Seg-diffusion outperforms previous well-established methods and achieves impressive zero-shot generalization to unseen datasets.

BibTeX
@inproceedings{icassp2025_segdiffusiontext,
  title = {Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation},
  author = {Shuo Zhang and Jiaming Huang and Yan Wu and Tao Hu and Wenbing Tang and Jing Liu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Seg-diffusion: Text-to-Image Diffusion Model for Open-Vocabulary Semantic Segmentation · ICASSP 2025