2025
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
AAAI 2025technical
Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image desc…