← Search

Xuanqi Chen

1 accepted papers

2026

ModalSyncSum: Synchronizing Image and Text for Reliable Summary Generation

AAAI 2026technical

Multimodal summarization with multimodal output (MSMO) aims to generate coherent textual summaries while selecting the most semantically relevant images to enhance expressiveness. Despite the advancements of large multimodal models like GPT-4o, LLaMA-3, and Grok-3, these models often exhibit halluci

Cited by 0SourcePDFScholar