Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
Multimodal keyphrase generation (MKP) aims to extract a concise set of keyphrases that capture the essential meaning of paired image–text inputs, enabling structured understanding, indexing, and retrieval of multimedia data across the web and social platforms. Success in this task demands effectivel