← Search

Soeun Lee

2 accepted papers

2025

ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning

AAAI 2025technical

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding. Because of this issue, there is a limitation that the image desc…

2024

IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning

EMNLP 2024main

Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To addre…