← Search

Yiqian Wu

4 accepted papers

2026

ModalSyncSum: Synchronizing Image and Text for Reliable Summary Generation

AAAI 2026technical

Multimodal summarization with multimodal output (MSMO) aims to generate coherent textual summaries while selecting the most semantically relevant images to enhance expressiveness. Despite the advancements of large multimodal models like GPT-4o, LLaMA-3, and Grok-3, these models often exhibit halluci

Cited by 0SourcePDFScholar
2025

EgoM2P: Egocentric Multimodal Multitask Pretraining

ICCV 2025accepted

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the camera wearer's actions, intentions, and surrounding environ…

Cited by 0SourcePDFScholar