2023
The CLIP Model is Secretly an Image-to-Prompt Converter
NeurIPS 2023poster
The Stable Diffusion model is a prominent text-to-image generation model that relies on a text prompt as its input, which is encoded using the Contrastive Language-Image Pre-Training (CLIP). However, text prompts have limitations when it comes to incorporating implicit information from reference ima…