2026
Knowledge-Enhanced Image Captioning with Adaptive Graph-based Multimodal Alignment and LLM
AAAI 2026technical
Image captioning is crucial for multimodal understanding, bridging visual content and natural language. Despite recent advancements in Large Multimodal Models (LMMs), when faced with unseen entities or scenes in the open world, even when attempting to leverage learned knowledge, models still struggl