Dual-Modality Guided Artistic Style Transfer with Pre-trained Diffusion Models
Jiaxiong Liu, Xiaolong Xiong, Jun Zhou
Abstract
Artistic style transfer aims to replicate an artist’s painting style in a different image. While existing pre-trained model-based methods can generate high-quality stylized images, they often lack precise control over stylistic elements. Recent approaches incorporating textual inversion offer more accurate style representations, but significant information loss occurs when transitioning between modalities. To address these issues, we propose Dual-Modality Guided Artistic Style Transfer (DMG), which makes full use of text and image information to enhance the visual effect and content consistency of stylized results. Our approach primarily consists of two key modules: Enhanced Style Encoding (ESE) and Guided Diffusion Generation (GDG). ESE processes information from both modalities to obtain an optimized and more comprehensive style representation. Subsequently, GDG employs stochastic inversion and attention control to ensure accurate delivery of content and style information. Our approach outperforms existing techniques in terms of visual quality and content consistency.
BibTeX
@inproceedings{icassp2025_dualmodalityguid,
title = {Dual-Modality Guided Artistic Style Transfer with Pre-trained Diffusion Models},
author = {Jiaxiong Liu and Xiaolong Xiong and Jun Zhou},
booktitle = {ICASSP 2025},
year = {2025}
}