AAAI 2026technical0 citations
FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation
Peng Zhang, Wanggui He, Mushui Liu, Wenyi Xiao, Siyu Zou, Yuan Li, Xingjian Wang, Guanghao Zhang
Abstract
Recent unified models have demonstrated that the reasoning capacity of Multimodal Large Language Models (MLLMs) can be leveraged to facilitate diffusion-based image generation with impressive flexibility and performance. However, approaches that rely heavily on MLLMs for high-level semantic encoding often struggle with fine-grained visual tasks like image editing and virtual try-on. To address this gap, we propose FUSE, a unified framework excelling at both high-level vision–language understanding and fine-grained generation. First, we introduce a Semantic-to-Detail Connector that pre-aligns fine-grained visual features with the MLLM
BibTeX
@inproceedings{aaai2026_fusefinegraineda,
title = {FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation},
author = {Peng Zhang and Wanggui He and Mushui Liu and Wenyi Xiao and Siyu Zou and Yuan Li and Xingjian Wang and Guanghao Zhang and Yanpeng Liu and Weilong Dai and Jinlong Liu and Shuyi Ying and Ruikai Zhou and Yunlong Yu and Yubo Tao and Hai Lin and Hao Jiang},
booktitle = {AAAI 2026},
year = {2026}
}