AAAI 2026technical0 citations

FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation

Peng Zhang, Wanggui He, Mushui Liu, Wenyi Xiao, Siyu Zou, Yuan Li, Xingjian Wang, Guanghao Zhang

Abstract

Recent unified models have demonstrated that the reasoning capacity of Multimodal Large Language Models (MLLMs) can be leveraged to facilitate diffusion-based image generation with impressive flexibility and performance. However, approaches that rely heavily on MLLMs for high-level semantic encoding often struggle with fine-grained visual tasks like image editing and virtual try-on. To address this gap, we propose FUSE, a unified framework excelling at both high-level vision–language understanding and fine-grained generation. First, we introduce a Semantic-to-Detail Connector that pre-aligns fine-grained visual features with the MLLM

BibTeX
@inproceedings{aaai2026_fusefinegraineda,
  title = {FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation},
  author = {Peng Zhang and Wanggui He and Mushui Liu and Wenyi Xiao and Siyu Zou and Yuan Li and Xingjian Wang and Guanghao Zhang and Yanpeng Liu and Weilong Dai and Jinlong Liu and Shuyi Ying and Ruikai Zhou and Yunlong Yu and Yubo Tao and Hai Lin and Hao Jiang},
  booktitle = {AAAI 2026},
  year = {2026}
}
FUSE: Fine-Grained and Semantic-Aware Learning for Unified Image Understanding and Generation · AAAI 2026