2025
Zero-Shot Image Captioning with Multi-type Entity Representations
AAAI 2025technical
As data and computational resources continue to expand, incorporating a variety of knowledge during the pre-training phase enhances large models, providing them with strong zero-shot capabilities. Due to the alignment of modal features by visual language models, zero-shot image captioning no longer…