Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query Refinement
GuangHao Meng, Jinpeng Wang, Qian-Wei Wang, XuDong Ren, Dan Zhao
Abstract
Vision-Language Retrieval (VLR) aims to retrieve relevant visual or textual information from multimodal data using language or image queries. However, traditional VLR methods often rely on data-driven shallow semantic alignment and fail to understand the deeper structural and fine-grained entity features of queries, resulting in poor performance on multi-entity layouts and challenging entities. In this paper, we propose the Layout-Aware and Sketch-Enhanced (LASE) VLR framework, which refines query representations by incorporating multimodal layout and sketch knowledge. Specifically, layout knowledge encodes the spatial arrangement of entities, while sketch knowledge refines entity perception by capturing essential structural details. To extract these knowledge representations, we leverage Large Language Models
BibTeX
@inproceedings{aaai2026_imaginewithlayou,
title = {Imagine with Layout and Sketch: Enhancing Vision-Language Retrieval with Dual-Stream Multi-Modal Query Refinement},
author = {GuangHao Meng and Jinpeng Wang and Qian-Wei Wang and XuDong Ren and Dan Zhao},
booktitle = {AAAI 2026},
year = {2026}
}