Knowledge-Enhanced Poetry-Image Synthesis with Large Language Model
Abstract
Classical Chinese poetry is a valuable cultural heritage. Transforming classical poetry into visual images enriches its expressive form and makes it more accessible to modern audiences. While generative models have succeeded in text-to-image synthesis, existing methods struggle with classical poetry due to data scarcity and semantic differences. These methods often fail to capture the imagery entities with ancient-modern semantic differences in classical poetry, resulting in deviations from the poem’s original intent. In this paper, we propose Poem-Diffusion, a novel knowledge-enhanced method that combines a latent diffusion model with a multimodal knowledge graph and transfers the knowledge from large language models to smaller models through knowledge distillation. As a result, the imagery embedded in classical poetry can be accurately converted into visual forms, even in resource-constrained environments.
BibTeX
@inproceedings{icassp2025_knowledgeenhance,
title = {Knowledge-Enhanced Poetry-Image Synthesis with Large Language Model},
author = {Simin Yang and Yuqing Li and Bin Wu},
booktitle = {ICASSP 2025},
year = {2025}
}