CVPR 20260 citations

WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation

Wei Chow, Jiachun Pan, Yongyuan Liang, Mingze Zhou, Xue Song, Liyu Jia, Saining Zhang, Siliang Tang

Abstract

Recent unified multimodal models (UMMs) have achieved remarkable progress in visual comprehension and generation. However, existing datasets and benchmarks focus predominantly on single-turn interactions, overlooking the multi-turn, context-dependent nature of real-world image creation and editing. To bridge this gap, we introduce W E A V E, the first comprehensive suite for in-context interleaved cross-modality comprehension and generation, comprising two complementary components. W E A V E-100k is a large-scale dataset of 100,000 interleaved samples spanning over 370,000 dialogue turns and 500,000 images, encompassing comprehension, editing, and generation tasks that demand reasoning over prior context. W E A V E-Bench is a human-annotated benchmark of 100 items with 480 images, equipped with a hybrid VLM judge evaluation framework that jointly leverages reference images and original-image-instruction pairs to assess multi-turn generation, visual memory, and world-knowledge reasoning across diverse domains. Experiments show that training on W E A V E-100k substantially improves vision comprehension, image editing, and comprehension-generation collaboration, while further enabling the emergence of visual-memory capabilities in UMMs. Extensive evaluations on W E A V E-Bench reveal persistent limitations of current approaches in multi-turn, context-aware image generation and editing. We hope W E A V E provides both a perspective and a foundation for advancing in-context interleaved comprehension and generation in the multimodal community.

BibTeX
@inproceedings{cvpr2026_weaveunleashinga,
  title = {WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation},
  author = {Wei Chow and Jiachun Pan and Yongyuan Liang and Mingze Zhou and Xue Song and Liyu Jia and Saining Zhang and Siliang Tang and Juncheng Li and Fengda Zhang and Weijia Wu and Hanwang Zhang and Tat-Seng Chua},
  booktitle = {CVPR 2026},
  year = {2026}
}
WEAVE: Unleashing and Benchmarking the In-context Interleaved Comprehension and Generation · CVPR 2026