← Search

Xiaoxing You

3 accepted papers

2026

Cut to the Chase: Training-free Multimodal Summarization via Chain-of-Events

CVPR 2026

Multimodal Summarization (MMS) aims to generate concise textual summaries by understanding and integrating information across videos, transcripts, and images. However, existing approaches still suffer from three main challenges: (1) reliance on domain-specific supervision, (2) implicit fusion with w

Cited by 0SourcecodeScholar
2026

Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning

AAAI 2026technical

News image captioning aims to produce journalistically informative descriptions by combining visual content with contextual cues from associated articles. Despite recent advances, existing methods struggle with three key challenges: (1) incomplete information coverage, (2) weak cross-modal alignment

Cited by 0SourcePDFScholar